shipwithjev

Catalog / Research & data

0568GitHub

Jev on BANKING77: 77-way support-intent classification

A frozen Jev Choice setup classifies 3,080 banking support messages using category definitions and retrieved labeled training examples.

Public API code, protocol, frozen implementation, and evaluation report were inspected on October 1, 2026. The author reports 2,846 correct classifications out of 3,080 test messages, or 92.40%, using 24 BM25-retrieved training examples per request. The cited 93.66% BERT result is a published comparison, not a newly run baseline. Inputs and examples were developed on training data and frozen before the test. This is one author-run dataset evaluation; labeled examples were still required.

simonmesmith/jev-banking77-experimentREADME ↗
# Jev on BANKING77: classification without task-specific fine-tuning

**An exploratory evaluation of a general-purpose decision model against a published specialist benchmark**  
September 18, 2026 · Model: `jev-1.13.0`

## Summary

Can a general-purpose model approach the accuracy of a classifier trained for one particular task, without being fine-tuned for that task itself?

We tested TypeSafe’s **Jev** on **BANKING77**, a public dataset of banking support messages with 77 intent categories. After comparing five input setups using training data only, we froze the strongest validation setup and evaluated all **3,080 official test messages**. Jev received category definitions and 24 relevant labeled training examples for each prediction; its weights were not updated.

**Jev classified 2,846 messages correctly: 92.40% accuracy.** The original BANKING77 paper reports **93.66%** for a fine-tuned BERT classifier, a gap of **1.26 percentage points**. Our final test cost an estimated **US$0.44** in API usage and took **6 minutes 51 seconds** with four concurrent workers. Development and testing together cost **US$0.77**.

This is encouraging evidence that Jev can provide useful classification without training a dedicated model. It is one public-benchmark result using labeled examples—not evidence that specialist classifiers are generally unnecessary or that Jev had never encountered similar data before.

## Background

### What is Jev?

[TypeSafe describes Jev](https://docs.typesafe.ai/concepts/system-one) as a “System One” model: an AI model designed to make structured decisions inside software. A developer provides information, a question and the allowed answers. Jev returns a typed judgment and probabilities rather than generating a free-form reply.

Its interface supports

Also filed under Research & data

  1. 0620

    Jev Score: rubrics, scores and confidence

    A worked guide to Jev's Score primitive: writing a request, defining rubric levels, reading recorded probabilities and the weighted-score math.

    Jev Trader · Research & data

  2. 0614

    Bot journey classification in WebDecoy

    Sends a detected bot's last 48 request paths to Jev, which picks what it's after (prices, articles…) and how it crawls (pagination, IDs…), or unknown.

    WebDecoy · Research & data

  3. 0610

    Fake / Real: link fact-checker with Jev as judge

    Paste an article or post link. Fake / Real extracts its claims, finds outside evidence, and has Jev judge whether the evidence supports or contradicts each one.

    @DansiDanutz · Research & data

  4. 0609

    AutoRubric: rubric-based evaluation

    Combines rubric science and LLM-as-a-judge research to grade outputs with AI judges: LLMs, decision models like Jev, or both.

    @deliprao · Research & data