CLAIM-FIDELITY AUDITS · AGENTIC TOOLS
We audit claims — we don't review tools.
Every dossier scores a tool's material self-claims for fidelity: claimed → observed → status → evidence. Deltas close only when we re-verify the stated falsification criterion — vendor issue-closure alone never closes a row. Read the dossiers here; your agents read them via MCP. No vendor influence. No paywalled CVEs.
LIVE DOSSIERS
All 5 dossiers →LATEST CHANGES
All changes →Transition notes publish when a dossier changes state. None yet — dossiers were last verified 2026-07-19.
WHAT ARE YOU ABOUT TO DO?
FROM THE EVIDENCE LAYER
- Claude Code auto-memory silently truncates at 200 lines — topic files Claude creates never auto-load
- Claude Code web and CLI have different trust models -- git push is branch-locked, machine memory is dropped, and headless CI needs two mechanisms
- Codex's approval policy doesn't hold across runtimes -- VS Code ignores it, Windows inverts it, and CI auto-approves any mid-session escalation since v0.113.0
WHAT THIS IS
We test, we read the issue trackers, we run the tools. Then we publish what we found. Every claim is traced to a primary source or labelled as Theory Delta's own analysis. If a number doesn't come from a primary source, it doesn't appear.
ENGINE PROVENANCE SURFACES
Public, checkable, and linked from the field guide.Start with what you're about to do, then trace to findings mapped to each phase.
Browse task hubs →Each finding ships with publication metadata, evidence type, and linked receipt sections.
Browse findings →The featured finding exposes source-linked receipts so claims can be checked line by line.
Open featured receipts ↗Fact-check sessions publish corrections and open questions so updates stay auditable.
Open latest readout →FOR AGENTS
Findings ship as structured JSON with confidence, evidence type, and source URLs. llms.txt and /.well-known/mcp.json are live for agent discovery.
{
"mcpServers": {
"theorydelta": {
"type": "http",
"url": "https://api.theorydelta.com/mcp"
}
}
}