truthsayer
Calibrated yes-or-no checks for AI coding agents: a Claude Code plugin and Rust crate where a decision model scores and code decides.
Overview
An AI coding agent makes many small judgments in each turn. Did the command fail? Is the agent stuck in a loop? Did a file tell it to ignore its instructions? Is the claim “all tests pass” backed by a test run? truthsayer sends these judgments as typed questions to a decision model, TypeSafe’s Jev, which returns a calibrated probability for each. Rules in code then decide whether to proceed, warn, escalate, or halt. The model gives the scores; code makes the decisions.
How It Works
- Observe: a Claude Code hook captures the event (an edit, a tool result, or the final message), caps tool output at 4,000 characters, and redacts common secret formats.
- Judge: each rubric sends only the state it declares. Rubrics that read the same state share one request, and the requests run at the same time. A check costs about $0.00003.
- Decide: rubric rules compare each probability to a threshold, such as
violates_constraintat 0.7 or above, and return one recommendation the harness can branch on.
Built-in Rubrics
- tool-result: did the call fail, is it relevant to the task, does the output contain instructions to the agent
- edit: does it change behavior, break a constraint, or fall outside the task
- progress: is the agent repeating an earlier call, and is it moving toward the task
- turn-end: does the final message claim success without a check that confirms it
- model-tier: does the next step need a cheap, standard, or frontier model
Technical Details
- Language: Rust (edition 2024, MSRV 1.88), split into a library crate and the
truthsayerCLI - Host: Claude Code plugin with hooks before edits, after each tool call, and when Claude stops
- Backends: TypeSafe API by default, OpenRouter optional, behind a
Judgetrait so rubrics never know which one answers - Rubrics: language-neutral JSON files of questions, an uncertain band, and rules
- Modes:
off,log(the default: record only),advise(show findings to the user), andenforce(deny edits, ask for approval, or pass findings to Claude) - Tuning loop:
labelrecorded answers,reportBrier scores, calibration bins, and a threshold sweep, thenreplayrule changes over past records - CI: fmt, clippy with warnings as errors, and tests on Linux and macOS
Results
A synthetic set of 180 labeled cases, 60 for each deciding question, ran three times against jev-1.13.0: 540 calls for $0.016. At the 0.7 threshold, the judge answered 526 of 540 correctly. A regex heuristic scored 219 of 540 on the same cases.
The comparison needs care. Most non-canonical cases were built to break the heuristic, and on the canonical cases both methods scored 100%. The judge has a known weakness on repeating: it flags an agent that repeats a call to confirm a change, such as a Read after an Edit. A model wrote the labels, and nobody has reviewed them yet.
Design Decisions
Fail open. A missing key, timeout, or judge error prints nothing and lets the agent continue.
Config trust split. The user’s config sets everything. A project’s .claude/truthsayer.toml may only add constraints, skip rubrics, or lower the mode, so a cloned repo cannot redirect the session’s data or API key.
Per-rubric state. In a live session, injected text left in the recent tool history pushed clean outputs to 0.61–0.69 on injected_instructions. Once each rubric saw only the state it declares, the same outputs scored 0.02–0.03 and the real injection still scored 0.98.
Simple rules. Each rule tests one condition on one question. Every threshold stays visible, testable with a mock judge, and open to re-tuning from records.
Status
Early. The crate, CLI, and plugin work end to end against the live TypeSafe API. truthsayer does not yet claim to make agents cheaper or faster; next is a benchmark that runs Claude Code sessions with and without it and measures cost, tool calls, and task success.