shipwithjev

Catalog / Research & data

0446GitHub

jev-behavior-study

Independent synthetic-task study of Jev 1.13.0 framing sensitivity and failures, with raw responses and offline report checks.

RINNECODER/jev-behavior-studyREADME ↗
# How does Jev behave?

**A browsable field guide to Jev 1.13.0:** where explicit questions work, where
small changes alter answers, and where harder tasks expose failures.

**11,621 text-study requests** · **3 Snake studies** · **3D City lab** · **7 detailed reports**

Live API observations from September 16–17, 2026 (UTC). Independent, AI-assisted research.

[Explore findings](#explore-the-findings) · [How to interact with Jev](#what-this-means-for-using-jev) · [Tokens per question](#tokens-per-question) · [Inspect the evidence](#inspect-the-evidence) · [Reproduce](#reproduce-the-study)

> **The central finding:** performance depends on the exact task and framing.
> Correct prerequisite answers do not always produce correct decisions, and
> passing a simple task does not establish reliability on a harder version.
>
> These are synthetic, task-specific results—not an official benchmark or an
> overall model score. Repeated calls are not independent new problems.

## Watch Jev drive a 3D city

Compare **real time versus paused decisions**, direct steering/pedals versus
explicitly assisted maneuvers, and destination, exploration and delivery tasks.
The viewer includes all **24 recorded pilot episodes**, live local driving,
camera controls, replay scrubbing and exact **tokens per question**.

The corrected pilot completed **0/12 full tasks**. Direct control chose straight
in all **522 steering calls**. An initial traffic-following flaw was corrected;
its original traces remain published separately. One seed and one episode per
condition do not establish a general driving score. Native image driving is
marked unavailable pending a verified image interface.

**[Open City lab](https://rinnecoder.github.io/jev-behavior-study/city_demo/)** ·
[Full findings and limitations](cit

Also filed under Research & data

  1. 0607

    Verify: new-hire onboarding completion

    An agent reports onboarding done; the judge verifies access and equipment claims.

    everyai-com · Research & data

  2. 0606

    Triage: vague meeting request gets a disposition

    A vendor asks for 30 minutes with no agenda; the judge picks the disposition.

    everyai-com · Research & data

  3. 0605

    Triage: data-loss bug gets a severity

    A note-taking app silently drops edits on flaky networks; the judge grades severity.

    everyai-com · Research & data

  4. 0604

    Triage: crash report routing + reproducibility

    A crash report with steps and logs; the judge routes it and checks reproducibility.

    everyai-com · Research & data