PriorBench: test Jev batching, confidence, and failure modes
A preregistered evaluation publishes request code, raw responses, and outcomes across 21 experiments on Jev decision behavior.
Public Jev client code, raw JSONL records, preregistration, and recount tooling were inspected on October 1, 2026. The project reports 5,721 calls and 50 predictions registered before collection, including falsified predictions. Experiments probe batching, option order, wording, confidence, concurrency, and out-of-scope inputs. It documents confident wrong decisions when no none-of-these option is offered. Its latency and threshold findings are specific to the tested version, gateway, and corpus; they are not universal deployment rules.
# PriorBench — Jev *Measured, pre-registered evaluation of AI systems.* **An independent, pre-registered evaluation of TypeSafe AI's Jev — 5,721 calls, 21 experiments, $0.176.** Model `typesafe/jev-1.13-20260917`, accessed through OpenRouter, measured from Western Europe on 20 September 2026. 📄 **[Read the full report →](paper/REPORT.md)** 📊 **[All 50 predictions and their outcomes →](paper/tables/predictions.md)** --- ## Headline findings - **Latency is a fixed cost, not a workload cost.** ~430 ms floor. **800 typed judgements in one call: 985 ms, $0.00075.** Chaining two Jev calls wastes 430 ms for nothing. - **95.9 % zero-shot** on our 400-item benchmark, versus 77.2 % for hand-written keywords and 66.0 % for supervised TF-IDF + logistic regression that saw every label. - **TypeSafe's own documentation understates the model.** Its "jaggedness" page says Jev cannot compare numbers and that date ordering is unreliable. We measure **99.6 % across 13 designs**. Negation, another listed failure mode, scores **100 %**. - **It always answers.** A cake recipe is classified as a technical issue at **0.94 confidence**; random letters at **0.97**. Without an explicit "none of these" option, **0 of 30** out-of-scope messages were flagged — at 0.99 confidence. - **Gate at 0.99 or not at all.** Accuracy above threshold is flat from 0.50 to 0.95, then jumps to **100 % at 0.99, covering 60.2 % of traffic**. - **Eight concurrent requests is the operating point.** Throughput saturates near 11 req/s. - **Wrong criteria descriptions are catastrophic:** 16.7 %, below the 25 % random floor. Missing descriptions cost only 0.8 points. ## This is the surface Everything here was produced in a single evening. It is a **wide, shallow pass**: 21 experimen