jev-lab
Experiments on TypeSafe's Jev decision model: measuring its repeatability and calibration, then building games where it plays Codenames and scores Wavelength clues.
Overview
Jev is a “System One” decision model from TypeSafe. It is not a chat model. You send a state and a map of typed questions, and it returns one calibrated answer per question: a yes probability, a pick among options, or a level on a rubric. jev-lab is a Bun workspace that tests those claims, then builds demos where code owns the workflow and Jev supplies narrow judgments.
Packages
- lab: a CLI that checks repeatability, sensitivity to rephrasing, and calibration through OpenRouter, with Claude Haiku 4.5 as a comparison arm
- codenames: a Codenames spymaster and guesser with a web app where every card shows Jev’s probability
- wavelength: a Wavelength round where code hides the target, Jev places every clue on the spectrum in one request, and the distance is the score
Findings
- Tight, not bitwise, repeatability. Over 20 repeats, the mean per-question spread in P(yes) was 0.002 to 0.017. A per-call nonce did not widen it, so the variance is inherent, not caching.
- Borderline answers still cross thresholds. One answer ranged 0.38 to 0.52 and flipped at a 0.5 cutoff 6 times in 20. The fix is an uncertain band from 0.3 to 0.7 that code routes elsewhere.
- Against Claude Haiku 4.5 at temperature 0. On the same rubric, Jev’s spread was 0.010 against Haiku’s 0.033. Jev ran 6× faster and cost about 55× less. On the three borderline questions, Haiku answered 0.85 or 0.15 with no spread; Jev sat near 0.5.
Codenames
Each clue is one request: 25 yes-or-no questions (“would a typical guesser connect this clue to this word?”) and one choice (“which word would they pick first?”). Code deals the board, checks clue legality, sets every threshold, and picks the clue. In one mode you type a clue and the board lights up in about 300 ms. In the other, Claude proposes candidate clues, Jev judges each against every word, and code picks.
On 300 human clues from the SALT-NLP Codenames Duet corpus, Jev matched the human’s first guess 69% of the time with a guess AUC of 0.935. Haiku 4.5 scored 60% and 0.880, at 1,674 ms and $0.0020 per call against Jev’s 227 ms and $0.00014. Haiku had the better Brier score, 0.047 against 0.066, so the data supports claims about ranking, speed, and cost, not calibration.
Technical Details
- Runtime: Bun workspace in TypeScript, one package per experiment
- API: the direct TypeSafe API through
@typesafe-ai/sdkfor the games; OpenRouter for the lab - Server:
Bun.servewith keys kept server-side, per-IP rate limits, and a streamed verdict feed over SSE - Replay: seeded boards, so any turn can be replayed, and a caching judge that replays answers from disk
- Tests: unit tests for the game engine and session state machine, with no network calls
What I Learned
The first side-by-side showed speed and cost but could not answer the obvious objection: a cheap chat model might do about as well. The fix was to run the baseline through the same eval as the model under test before writing any claim, and to publish the results that favor the baseline too.