Introducing XonTools: a hypothesis, six documents, and every result so far
What this site is, what the research program is testing, what has worked, what hasn't, and what comes next.
AI systems are getting more capable faster than our ability to check what they’re doing. This site is where I’m working on one piece of that problem in the open: tools that show where an AI system’s account of its work, or of itself, doesn’t fit the evidence.
This first post is a map: what the program is betting on, what’s here, what has been tested so far, and what comes next.
The bet
The research program is organized around one idea, the Extreme Oversight Hypothesis. Imagine a large, diverse, carefully validated set of windows: small read-only probes, each attached to a model’s internal states and reporting one thing, such as whether it is straining against its sources, guessing, or pursuing a goal other than the one it states. The hypothesis is that past some scale of oversight, being honest becomes cheaper for training to find than evading every window at once. An honest model has nothing to hide. A deceptive one has to hide everything, from every window, at the same time.
It may well be wrong, and the write-up spends as much time on how it could fail as on why it might hold. The hardest problem is at its center: a window used as a training signal learns to be fooled, while a window never used exerts no pressure at all. The two versions of the hypothesis answer that tension differently, and both can be probed with small experiments now.
Six documents
The ideas are set out in a numbered series, meant to be read in order: the claim, the instruments, then the architecture.
- The Extreme Oversight Hypothesis. The claim, its two versions, and every way it could be wrong.
- Windows into the Black Box. Thirty candidate windows in five families: how each would be built, what would teach it, and when in training it could first be used.
- XonTools: a vision of the finished system. What the tools are for, and the path to a released version.
- The Strain Head. The first internal window as a concrete, budgeted experiment: does an ordinary language model carry a readable signal of where it is being inconsistent?
- The Xon Neural Network. A network that computes by settling around fixed evidence, and the tests it would have to pass.
- Hybrid Minds. How settling modules, language models and exact solvers could be composed into one system that can say where it strains.
All six are on the research page, free to read and reuse under CC BY 4.0.
What has been tested
The first instrument is a consistency engine: it turns a document into claims and the relations between them, and finds contradictions exactly, including ones that no pair of claims reveals on its own. On its first pre-registered benchmark, 180 synthetic documents, it found planted cycles of jointly impossible claims with an F1 of 0.92, where checking claims two at a time managed only 0.33.
It’s a good first result, and it will get better. The weak spot is precision: the engine wrongly flagged 18% of consistent documents, and a strong language model asked directly scored slightly higher, 0.93, at about a twentieth of the cost. So the engine’s case today is not raw accuracy but structure: every flag comes with the specific claims and relations behind it, which a person can check.
The false positives have already been traced to their source. Ten of the eleven came from the step that judges how two claims relate, which was applying a too-loose definition of “contradicts,” and the eleventh from a comparison read in the wrong direction. None was a real contradiction hiding in the text. A precision revision fixes both causes: a strict definition of contradiction, each pair of claims judged in isolation, and directions derived by code rather than by the model. On the original benchmark it cut the false-positive rate from 18% to about 7% and raised F1 from 0.92 to 0.97, without losing a single detection. It also roughly tripled the cost, to about $0.42 per document, because judging every pair separately takes more model calls. At this stage that’s a price worth paying: a few hundred dollars answers the question that matters, whether structure buys something a single judge can’t, and cost is something to engineer down once the answer is known. The improved accuracy figures come from documents the engine has already seen, though, so they guide development and don’t count as a result. The test that does count is pending: a single run on a fresh, sealed test set, with its commitment posted before it starts. I expect it to come in lower than the development numbers, and it will be published whatever it shows.
Not everything worked. The idea this project began with, that a rich “harmony” in oscillator fields could serve as a goal for safe systems, failed its first tests in three different ways, including one where the metric rose under exactly the kind of drift it was supposed to detect. What survived was sharper: consistency is worth measuring only against evidence the system isn’t allowed to change. The Origins page tells that story.
Everything, including what failed
The results page lists every run so far: passes, failures, mixed outcomes and work in development, each with its date, the criteria that were fixed before it ran, and links to the data and code. Each result also carries a provenance level, so you can see how much to trust it:
- Scouting: a quick test at toy scale. A reason to design a real experiment, not a result.
- Recorded: criteria written down before the run, and every model call cached so the run can be replayed exactly.
- Committed: all of that, plus a timestamped public commitment, the frozen code and a hash of the sealed test set, posted before the run, so nobody can move the goalposts afterwards.
The first committed run, the engine’s precision revision, will have its commitment posted there before it starts.
Open by default
The code is open source under Apache 2.0 at github.com/IboTool/XonTools, and the documents and site content are under CC BY 4.0. The only things kept private are the few that only work while hidden, such as fresh test sets, and those are committed in public by their hashes before they’re used. The Origins page sets out the commitments in full.
What comes next
- The engine’s precision revision, run once on a fresh sealed corpus, with its commitment posted first.
- XonForge, a generator for larger, harder test sets whose ground truth comes from a solver rather than from any model.
- The agent transcript monitor, which checks what an AI agent reports against what its tools actually returned. It’s the part of this work that matters most in practice.
- The strain head, the first window inside a model, with its questions and stop rules published before it runs.
How to help
The most useful thing anyone can do right now is try to break the ideas. If you think the hypothesis fails in a way the write-up doesn’t cover, or you know of work that already tests it, I’d like to hear it. If you work on oversight, interpretability or agent evaluation and want to compare notes or share test sets, get in touch. And if you fund independent safety research, grant applications are under way, and I’d be glad to talk.
— Ian