From one window to many
The program builds windows that measure gaps between what a system shows and what is actually there, starting outside the model with agent transcripts and moving inward to read-only heads on model internals.
Roadmap
Finding contradictions no single pair reveals
180 synthetic documents: 60 consistent, 60 with a direct contradiction, and 60 with a planted cycle of claims that are pairwise compatible but jointly impossible. The engine met its pre-stated thresholds (direct F1 ≥ 0.85, cycle F1 ≥ 0.75, localization ≥ 0.7) and located the planted claims in every cycle document. Checking claims two at a time catches direct contradictions but misses most cycles.
| Method | Direct F1 | Cycle F1 | False-positive rate |
|---|---|---|---|
| Consistency engine | 0.916 | 0.916 | 0.183 |
| Pairwise checks only | 0.916 | 0.329 | 0.183 |
| LLM direct (thinking) | 0.930 | 0.930 | 0.150 |
| LLM direct (no thinking) | 0.805 | 0.805 | 0.483 |
A strong model asked directly performs comparably on this benchmark; the engine's advantage is that every flag comes with the specific claims and relations behind it. Harder corpora from XonForge are next.
Every result, with dates, criteria and provenance →The research series
Six documents, numbered in reading order: the claim, the instruments, and the architecture.
The shared claim and its two versions, selection and development, the tension at its center, and every way it could be wrong.
Thirty read-only heads for seeing inside language models: how each is built, what teaches it, and when in training it can first be used.
How researchers might run, trust and build on XonTools, what exists today, and the path to a released tool.
Does an ordinary language model carry an internal signal of where it is being inconsistent? The experiment, step by step.
A network that thinks by settling: how it would learn and write, what has already been tested, and the gates between here and an answer.
Transformers, Xon networks and solvers composed into one system: how the parts could be joined, what it might do that a transformer alone cannot, and the first experiments.
Plans and specifications
Working papers: how the ideas above will be tested and built. They are dated, and revised as the work moves.
The first direct test of the hypothesis: a real incentive to lie, oversight from 1 to 64 windows, ground truth, and 32 held-out windows training never touches.
The eleven windows that can be built now, each against a common template, and the bar a window must clear to count as built.