shipwithjev

Blog / Builds & people / FIG. 138

Jev, One Month In: What Held Up

Jev month one so far: 551 cataloged builds, the claims that held up, the ones that cracked, and what the ecosystem looks like heading into October.

TypeSafe AI opened access to Jev, its decision model, on September 15. Month one isn't quite over as we write this at the end of September, so treat this as "Jev month one, so far": a check on which launch-week claims survived another two weeks of builders poking at them. The week-one story itself (speed proofs, economics proofs, the cascade emerging) is told in the launch week recap; this page picks up where that one stops.

Everything below comes from the directory's own entries. Numbers are as reported by builders and linked. No official benchmarks exist, and nothing here should be read as one.

The updated counts

The catalog holds 551 entries at the time of writing. Filing dates run from September 15 to September 23, so the most recent week isn't reflected yet; expect the count to move when the next batch is sorted.

  • By category: tools and apps 213, agents and browsers 79, research and data 72, games and real-time 66, content and growth 54, triage and routing 50, trading and markets 10, robotics and devices 7.
  • By source: 278 GitHub repositories, 183 X posts, 46 write-ups and resources, 26 Reddit posts, 9 sites, 9 agent skills.
  • With numbers attached: 132 entries carry at least one reported figure (cost, latency or a stat), and 45 report a cost.

The shift worth noticing is in the source mix. Before September 20, entries were mostly posts and demos. From the 20th on, the bulk was repositories, and tools and apps went from 42 entries to 213. The ecosystem moved from showing Jev off to building things around it.

What held up

The cost pattern. Later receipts kept landing where early ones did. 100,000 posts for $0.67, 22.8 million log lines for $0.64, 6,245 companies for $0.37, all as reported. Nobody has published a decision-shaped workload that contradicts the launch-week economics.

The cascade as the default shape. The fraud cascade was a clever build in week one. By week two there were routers, gatekeepers and pick-then-write pipelines all over the catalog, compared in five cascade architectures.

Real deployments, not only demos. Metaview reports shipping Jev into every agent on its platform, with candidate searches going from minutes to seconds at the same accuracy. AgentMail classifies every outbound email before sending. Ryze reports a 90% cost cut on its SEO agents. All self-reported; all worth watching.

What cracked, or never claimed to hold

Fast is not strong. The LLM Chess row is 8 wins and 50 losses in 80 games against a random opponent, with zero illegal moves, as reported.

Fast is not always fast enough. Ms. Pac-Man measured about 650 ms average against an estimated 70 ms needed for real-time play.

Jev alone is not the whole agent. On the WebMCP benchmark, Jev driving the page directly solved 25 of 49 tasks; paired with a model that wrote the arguments, 49 of 49. The split works. The solo act doesn't.

None of these contradict what TypeSafe said Jev is. They contradict what the hype said it was, and the limitations page predicted most of them.

What's new since launch week

Integrations arrived. Adapters or providers for LangChain, Pydantic AI, TanStack AI, Effect and LiteLLM, plus community clients for Go, .NET, Rust, R, Elixir, Swift, Laravel and Spring Boot. Check each against docs.typesafe.ai; community packages are not official.

Coding-agent guardrails became a genre. jev-gates, Bouncer at a reported ~$0.04 a day, a shell-command safety hook and more. The shared rule: escalate or block on a verdict, never auto-approve something irreversible on one.

Independent evaluations started. jev-evaluation fixed 28 predictions before running 123,805 requests for $12.69, as reported, and publishes what held and what didn't. jev-decision-benchmarks tests tool selection and abstention. Community evaluations, not official benchmarks, but real ones.

Our own verdict from the Jev review stands: exceptional at one thing, useless at the rest, honest about the split. Month one so far has only sharpened the edges.

Frequently asked questions

How many Jev builds exist after one month?

The directory catalogs 551 as of the end of September, with filing dates through September 23. That is a floor, not a census.

Did Jev's cost claims hold up?

For decision-shaped workloads, yes: later builder receipts, from $0.37 to $0.67 for tens of thousands of decisions, match the launch-week pattern. See what builds actually cost.

Are there official Jev benchmarks yet?

Not that we can cite. Community evaluations like jev-evaluation and jev-decision-benchmarks exist and are linked above, but they are independent work, not TypeSafe's.

What changed most since launch week?

The source mix: from X demos to GitHub repositories, integrations and guardrail tools. The launch week recap is the before picture.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.