shipwithjev

Catalog / Research & data

0569GitHub

PriorBench: test Jev batching, confidence, and failure modes

A preregistered evaluation publishes request code, raw responses, and outcomes across 21 experiments on Jev decision behavior.

Public Jev client code, raw JSONL records, preregistration, and recount tooling were inspected on October 1, 2026. The project reports 5,721 calls and 50 predictions registered before collection, including falsified predictions. Experiments probe batching, option order, wording, confidence, concurrency, and out-of-scope inputs. It documents confident wrong decisions when no none-of-these option is offered. Its latency and threshold findings are specific to the tested version, gateway, and corpus; they are not universal deployment rules.

priorbench/jevREADME ↗
# PriorBench — Jev

*Measured, pre-registered evaluation of AI systems.*

**An independent, pre-registered evaluation of TypeSafe AI's Jev — 5,721 calls,
21 experiments, $0.176.**

Model `typesafe/jev-1.13-20260917`, accessed through OpenRouter, measured from
Western Europe on 20 September 2026.

📄 **[Read the full report →](paper/REPORT.md)**
📊 **[All 50 predictions and their outcomes →](paper/tables/predictions.md)**



---

## Headline findings

- **Latency is a fixed cost, not a workload cost.** ~430 ms floor. **800 typed
  judgements in one call: 985 ms, $0.00075.** Chaining two Jev calls wastes
  430 ms for nothing.
- **95.9 % zero-shot** on our 400-item benchmark, versus 77.2 % for hand-written
  keywords and 66.0 % for supervised TF-IDF + logistic regression that saw every
  label.
- **TypeSafe's own documentation understates the model.** Its "jaggedness" page
  says Jev cannot compare numbers and that date ordering is unreliable. We
  measure **99.6 % across 13 designs**. Negation, another listed failure mode,
  scores **100 %**.
- **It always answers.** A cake recipe is classified as a technical issue at
  **0.94 confidence**; random letters at **0.97**. Without an explicit "none of
  these" option, **0 of 30** out-of-scope messages were flagged — at 0.99
  confidence.


- **Gate at 0.99 or not at all.** Accuracy above threshold is flat from 0.50 to
  0.95, then jumps to **100 % at 0.99, covering 60.2 % of traffic**.
- **Eight concurrent requests is the operating point.** Throughput saturates near
  11 req/s.
- **Wrong criteria descriptions are catastrophic:** 16.7 %, below the 25 % random
  floor. Missing descriptions cost only 0.8 points.

## This is the surface

Everything here was produced in a single evening. It is a **wide, shallow
pass**: 21 experimen

Also filed under Research & data

  1. 0620

    Jev Score: rubrics, scores and confidence

    A worked guide to Jev's Score primitive: writing a request, defining rubric levels, reading recorded probabilities and the weighted-score math.

    Jev Trader · Research & data

  2. 0614

    Bot journey classification in WebDecoy

    Sends a detected bot's last 48 request paths to Jev, which picks what it's after (prices, articles…) and how it crawls (pagination, IDs…), or unknown.

    WebDecoy · Research & data

  3. 0610

    Fake / Real: link fact-checker with Jev as judge

    Paste an article or post link. Fake / Real extracts its claims, finds outside evidence, and has Jev judge whether the evidence supports or contradicts each one.

    @DansiDanutz · Research & data

  4. 0609

    AutoRubric: rubric-based evaluation

    Combines rubric science and LLM-as-a-judge research to grade outputs with AI judges: LLMs, decision models like Jev, or both.

    @deliprao · Research & data