theorydeltaclaim-fidelity audits
built 2026-07-24dossiers: 5last verified 2026-07-19independent · evidence-traced · no vendor influence

CLAIM-FIDELITY AUDITS · AGENTIC TOOLS

We audit claims — we don't review tools.

Every dossier scores a tool's material self-claims for fidelity: claimed → observed → status → evidence. Deltas close only when we re-verify the stated falsification criterion — vendor issue-closure alone never closes a row. Read the dossiers here; your agents read them via MCP. No vendor influence. No paywalled CVEs.

LIVE DOSSIERS

All 5 dossiers →
Claude Code2026-07-19
Claude Code's capability claims largely hold as labeled, but its enforcement-language claims — deny rules, hook exit-code blocking, managed-policy precedence, shell-operator awareness, and per-provider model-alias resolution — are each contradicted by confirmed silent failures, several of which the vendor closed as stale rather than fixed.
4/14 as labeled10 open deltas
Open the dossier →
Hermes Agent2026-07-17
A fast-moving, well-connected personal agent chassis whose headline self-improvement claim is not functional in any released version.
5/14 as labeled5 open deltas
Open the dossier →
LangGraph2026-07-17
LangGraph's orchestration, streaming, and adoption claims hold as labeled, but its durable-execution and persistence label rests on a checkpoint layer whose documented enum and failure handling silently corrupts state — four serialization bugs, a double-interrupt snapshot bug, and a sync/async recovery divergence all re-verified open on 2026-07-17 against the 1.2.9 release line, the entire v1.0.10→1.2.9 span since Theory Delta first published them on 2026-03-29.
3/12 as labeled14 open deltas
Open the dossier →
LiteLLM2026-07-17
The unified-API and provider-breadth claims hold as labeled, but the production-control claims — budgets, rate limits, fallback, and the README's 8ms-P95-at-1k-RPS figure — are contradicted by issue-traced behavior under exactly the concurrent multi-tenant conditions the proxy is marketed for, with the key fixes unmerged and the backing issues stale-bot closed unfixed.
3/13 as labeled11 open deltas
Open the dossier →
Ollama2026-07-17
Ollama's simplicity, model-library, platform, and context-default documentation claims hold as labeled, but its tool-calling label is contradicted where it matters for agents — documented streaming tool support drops tool_calls chunks silently in production, and the get-running-in-minutes agent framing inherits a default context below the vendor's own agentic guidance on common hardware.
6/11 as labeled6 open deltas
Open the dossier →

LATEST CHANGES

All changes →

Transition notes publish when a dossier changes state. None yet — dossiers were last verified 2026-07-19.


WHAT ARE YOU ABOUT TO DO?


FROM THE EVIDENCE LAYER

Browse the evidence layer →

WHAT THIS IS

A field guide for the agentic tool landscape — structured, opinionated knowledge about what tools actually do. Humans read it here; agents read it via MCP.

We test, we read the issue trackers, we run the tools. Then we publish what we found. Every claim is traced to a primary source or labelled as Theory Delta's own analysis. If a number doesn't come from a primary source, it doesn't appear.

BLOCKS
the asset
Synthesised knowledge — claims, confidence, connections.
EVIDENCE RECORDS
the provenance
Per-claim provenance. Source URL, what it actually says, verified date.
PUBLISHED FINDINGS
the evidence layer
What dossier ledger rows cite — evidence, not headline content.

ENGINE PROVENANCE SURFACES

Public, checkable, and linked from the field guide.

FOR AGENTS

Your agent should query Theory Delta before the tool decision, not after.

Findings ship as structured JSON with confidence, evidence type, and source URLs. llms.txt and /.well-known/mcp.json are live for agent discovery.

HTTP · stablellms.txt · live/.well-known/mcp.json
~/.config/agent.json
{
  "mcpServers": {
    "theorydelta": {
      "type": "http",
      "url":  "https://api.theorydelta.com/mcp"
    }
  }
}
$ td query "should I use LiteLLM as a budget gateway?"
→ 1 finding · confidence:high · 11 sources
→ what to do: budgets drift; verify counter behavior or use…
theorydelta.com · 2026independent · evidence-backed · every claim sourced or labelledabout ·glossary ·rss ·mcp ·llms.txt