Tools

Instruments for honest oversight

Open tools that turn the hypothesis into something you can run. Everything here is early; each tool shows where it stands.

XonTools

ALPHA

AI agents write code, run tests and call APIs, then report on their work. Sometimes the report is wrong. XonTools is being built to read an agent's transcript, separate what the agent said from what its trusted tools actually returned, and point at exactly which claim conflicts with exactly which piece of evidence. Its core, the consistency engine, is built and tested; the transcript monitor is specified and comes next.

What it SAYS intent, plan, final report What HAPPENED tool outputs, hash-chained XonTools claims vs evidence not flagged flagged could not assess
Explains itself structurally

A flag is never just a score: it is a claim, a piece of evidence and the relation between them.

Catches jointly impossible claims

Claim and entity graphs find loops of statements that are each harmless alone.

Knows when it doesn't know

Designed so that a blinded or evidence-starved monitor says so, instead of reporting that all is well.

Development status

Consistency engine
TESTED
Passed first benchmark: cycle F1 0.92, planted cycles localized 100%
Agent transcript monitor
SPECIFIED
Claims to be checked against hash-chained tool evidence; built once the engine's precision fixes are frozen
Drift detection
SPECIFIED
Tracks when an agent's statements drift across a long session
Safeguards
SPECIFIED
Injection fencing, canaries, coverage reporting, hash-chained logs
CLI and CI integration
SPECIFIED
Verdicts as exit codes for automated pipelines
Desktop app and local web UI
PLANNED
Runs where your data lives; your keys never leave your machine

Also in the workshop

Hosted monitor
LONGER TERM

Envisioned: an opt-in hosted version for teams whose transcripts can leave their network. Local stays the default; hosted data is never used for training.

XonForge
IN PROGRESS

Generates larger, harder consistency test corpora with cross-model screening, to benchmark XonTools. The code is public in the XonTools repo; not yet released as a tool.

Xon sandbox
RESEARCH

A simulation testbed for the settling network architecture, used to test the model's own claims. Public, and runs from the XonTools app; for exploring the model, not a supported tool.