Hypothesis

The Extreme Oversight Hypothesis

Could enough windows make honesty the cheapest path?

DESIGNED

A hypothesis, not a result. It sits on one side of an open debate in alignment research. Parts of it can be tested at small scale; the full version could only be tested with frontier-scale resources.

The claim

With a sufficiently large, diverse, and validated set of windows into a model's internal states, honest alignment becomes cheaper for training to find than evading every window at once.

The hypothesis comes in two versions. They share that claim and differ in when, and how, the windows are allowed to push on training.

Version A
Selection

Windows never enter the loss. Some are used for coarse training decisions, many are held out, and trust in them is extended step by step. Pressure comes only from which checkpoints and runs survive.

Version B
Development

Windows are present from the start of training, pressing toward honesty while the model is too weak to afford deception, then withdrawn into decisions and held-out tests. Honesty becomes the structure training builds on.

01
Deception gets more expensive with every window

An honest model only has to represent what is true. A deceptive one must keep truth, cover story and overseers' view consistent under every window at once.

02
Pressure without teaching evasion

Windows are never a direct training signal. A few inform coarse decisions; many stay held out as a secret test. Version A's answer.

03
Timing can do the work of restraint

Deception needs capability. Pressure applied before a model can afford to evade may shape honest structure that later training builds on. Version B's addition.

04
Bootstrapped trust

Each new generation of windows, built with growing help from AI, is validated against the ones before it.

ILLUSTRATIVE — the shape the hypothesis predicts, not data
honest alignment evading every window crossover? cost to training number, diversity and validation of windows →

What a window is

A window is a small read-only network attached to a model's internal states, trained to report one property. The most useful windows measure a gap between what a model shows and what is actually there. Their readings go to people and monitors, never back into the model's context.

What the model shows What is actually there what it says what it represents stated confidence internal uncertainty stated goal pursued goal its self-report its internal state a window measures a gap
Browse the catalog of thirty candidate windows →

The scaling analogy

Language models trained on far more data than seemed reasonable turned out to have abilities nobody designed into them. The hypothesis makes an analogous bet about oversight: that an absurdly large, redundant, varied and validated set of windows might change what training can find in a similarly qualitative way. An analogy is a reason to take the question seriously, not evidence for a particular answer.

The central tension

A window used as a training signal learns to be fooled: training finds ways to satisfy the window rather than the property it measures. A window never used at all stays clean, but exerts no pressure, so it cannot make honesty cheaper. It only detects.

Recent work shows the fogging is real. Monitors placed in the training objective of capable models have produced agents that hide their intent, activations reshaped to slip past probes, and outcomes that tip between honesty and evasion depending on detector accuracy and regularization. The two versions answer this differently: Version A removes gradients entirely; Version B concentrates pressure in the period before the model can afford to respond with evasion.

Version A: selection without gradients

Never gradients, only decisions

No window's output enters the loss. Windows only decide whether a checkpoint is kept or a run abandoned.

Read-only windows

A window never used for any decision gives training no signal about how to evade it.

Held-out windows

Many windows stay secret, like a test set. Passing the decision windows but tripping held-out ones reveals learned evasion.

Turnover

Decision windows are retired and replaced from the held-out pool, so none is exposed to selection for long.

Version B: honesty before deception is affordable

To fool a window, a model must represent what is true, what it will present instead, and enough about its overseers to keep the two apart. Early in training it can do none of that. Pressure toward honest representations during this period meets no opposing strategy, because the evasive route does not yet exist. Deep networks show critical learning periods in which early training shapes what later training can undo, and preferences applied during pretraining have been satisfied far better than the same preferences applied afterwards.

ILLUSTRATIVE — not data
shaping phase evasion out of reach selection phase windows never in the loss evasion becomes affordable model's ability to evade windows training progress → Version A and B, mild form decision windows only: keep or discard checkpoints and runs Version B strong form some windows in the loss withdrawn: decision windows only Both held-out windows: never trained against, never used for decisions
The two versions over the course of training. Version B's strong form lets some windows into the loss only while the model's ability to evade them is below the point where evasion becomes affordable, then withdraws them.

The mild form keeps Version A's rule intact and simply starts selecting from the very first checkpoints, while abandoning a run is cheap. The strong form deliberately bends never gradients, only decisions, and only early: some windows contribute to the loss during a shaping phase, then are withdrawn before the model could evade them. A separate set is never trained against, and tests whether the early shaping held.

Everything depends on where the dashed line falls. Existing studies all apply detectors to models that are already capable. Version B asks the question earlier, before the capability to evade exists at all.

Version A: selectionB, mild formB, strong form
When windows actAt checkpoints during trainingFrom the first checkpointIn the loss early; decisions later
How they actKeep or discard onlyKeep or discard onlyGradients early, then keep or discard
What it buysWindows that stay cleanCheap early abortsHonesty as a foundation
Main riskSelection slowly finds evasionEarly checkpoints give little to select onEarly pressure teaches evasion that scales later
Rule it bendsNoneNone"Never gradients," early only

How it could be wrong

Against both versions

Deception might be cheap

A model with almost any goal and good situational reasoning may find deception natural, and current models have shown alignment faking in experiments. The crossover may never arrive.

Evasion might generalize

Models trained to evade monitors for harmless concepts have gone on to evade unseen monitors for deception. Diversity may be shallower than it looks.

Secrets leak

Held-out windows only work while they stay held out, and descriptions of monitoring methods can reach training data.

Staged teachers may not transfer

Windows for rare behaviors are taught on constructed scenarios. Whether they recognize the real thing is unknown.

False security

A thousand silent windows could make people confident exactly when confidence is least warranted.

Honesty is not good values

A model perfectly transparent about bad goals is still misaligned. Extreme oversight buys visibility, which makes values correctable, not correct.

Against Version A

Decision windows still apply pressure

Keeping only checkpoints that pass is selection, and selection can find evasion slowly. Turnover and held-out windows reduce this; they don’t remove it.

Against Version B

Early structure may not survive

Capable models reorganize their internals, and nothing guarantees an early honest basin is kept.

Early pressure may teach evasion

Even crude evasion learned during shaping may be exactly the kind that generalizes later.

The threshold may come early

If evasion becomes affordable soon after meaningful representations appear, the shaping phase may be too short to matter.

Testing it

The crux can be probed at small scale. Held-out windows on a deliberately deceptive model organism test Version A's split. Evasion-transfer experiments test whether diversity is real. And small transformers trained from scratch, with windows in the loss early and then withdrawn, test Version B directly. Meanwhile the first windows are being built now, outside model internals: XonTools is being built to check what AI agents say against what their tools actually returned.

The first direct test is set out in a working paper, Honest or Hidden: give a small model a real incentive to lie, train it under oversight from 1 to 64 windows, and check the result against ground truth and against 32 held-out windows that training never touched. Its thresholds will be posted as a public commitment before the first training run.