shipwithjev

Blog / Builds & people / FIG. 93

Jev for QA: The Checks Assertions Can't Write

AI software testing for QA teams: semantic test checks that assertions can't express, reported E2E run costs, and where Jev stays out.

Every QA engineer keeps a private list of bugs no assertion could have caught: the error message that was present and useless, the empty state that rendered fine and said nothing, the test that passed because it checked the wrong thing. That list is where AI software testing actually earns its keep. Jev, the decision model from TypeSafe AI, doesn't write your tests or invent test data; it answers closed questions about what your system produced, fast and cheap enough to run on every build.

The developer-side view of CI gates lives in Jev for developers. This page is the QA seat: coverage, flake, and the judgment layer between "passed" and "correct".

Where assertions run out and AI software testing starts

Assertions are rulings on exact values. They're perfect for "status is 200" and useless for "this error message tells the user what to do next." QA teams have covered that gap with manual passes, screenshot diffs a human eyeballs, and a certain amount of hope.

Semantic test checks turn the mushy middle into closed questions. Does this error message name the field that failed? Does the confirmation page state the amount the user entered? Is this empty state an explanation or a blank? Per ecosystem documentation, each answer comes back as a probability over the choices you define (yes, no, unclear by default), which means it can be thresholded, logged, and trended like any other test result.

The questions are the test cases. Write them the way you'd write acceptance criteria: specific enough that two testers reading the same output would agree.

The tests that lie

The sneakiest QA bug is a green test that verifies nothing. One cataloged tool, jev-lint, asks the questions a senior reviewer asks: does this function do what its name says, is this comment still true, does this test verify what it claims, with a cutoff per rule (behavior as described by its author). Pointed at a test suite, that's a coverage audit that reads intent instead of counting lines.

Same family, stranger idea: jevcumber runs Cucumber scenarios from the .feature file alone, with no step definitions, per its author. Whether that holds for your suite is an experiment, not a promise. It does show where the Gherkin-to-glue tax might be headed.

E2E runs where cost used to be the blocker

Agent-driven end-to-end testing has been stuck on economics. A frontier model stepping through a checkout flow is slow and pricey per run, so it runs nightly at best, and nightly means the regression already merged.

The receipt that changes the napkin math is jev-e2e, which writes tests in plain English and runs them with Jev and Playwright. On the same eBay flow, the author reports completed-run medians of 47 seconds and $0.0067 for Jev, against 62 seconds and $0.0277 for GPT-5.6 Luna and 79 seconds and $0.4062 for Claude Sonnet 5. That's as reported, one flow, one builder. There are no official Jev benchmarks, so read it as a hint that per-PR agent E2E is leaving budget-meeting territory, not as a guarantee.

Another project, e2e-testing-agents, is building an open-source framework for agent-run tests across web and mobile with Jev in the loop. Early, but the direction is consistent: an agent's step decisions (which element, is this the expected screen, did the action land) are closed-set, and closed-set is where a decision model is cheap.

Flake triage without the 2 a.m. guesswork

QA owns the flake budget, and flake triage decomposes into verdicts. Does this failure match a known flaky signature? Is the error network-shaped or logic-shaped? Did the failing test touch code in this diff? When the failure shows up in production instead of CI, the same verdict shape drives AI incident triage: real or noise, how severe, which team owns it.

Route the confident "known flake" slice to a quarantine lane, the confident "real regression" slice to the owning team, and the unclear middle to a person. The QA-specific rule is simple: a verdict can label, route, and page, but it never auto-deletes a test or auto-closes a bug on its own. Quarantine is reversible. Deletion is how a real regression gets a permanent hiding spot.

Treat the judge like code under test

A semantic check is itself a test that needs tests. Build a small gold set (a few dozen known-good and known-bad outputs per check), measure how often the verdict agrees with you, and rerun it whenever you change the question wording, exactly as the evals playbook prescribes. Version questions next to the tests that use them.

Know the edges, too. Jev does not generate text. If you need test data, bug-report prose, or repro steps written, that's a chat model's job; Jev scores what the chat model wrote. For visual regressions, check whether it judges images before designing around it.

A check you haven't measured is a vibe with a green checkmark.

Frequently asked questions

Can AI replace QA engineers?

No. It replaces the repetitive reading part of QA, the "does this output make sense" pass over thousands of results. Deciding what to test, building gold sets, and owning release calls stay with people.

What are semantic test checks?

Closed questions about system output that an exact-match assertion can't express, such as "does this error message name the invalid field?" Each returns a probability you threshold like any other result, per the evals approach.

Is AI E2E testing cheap enough for every pull request?

One builder reports 47 seconds and $0.0067 per completed run on a single eBay flow (jev-e2e), as reported and not independently verified. Measure your own flows before you budget.

Will verdict-based checks make my suite flaky?

They can if the question is vague or the threshold sits on the fence. Use three outcomes (pass, fail, unclear goes to a human), measure agreement against a gold set, and treat any flapping check as a question-design bug.

Where should a QA team start?

Pick one check your team already does by hand that no assertion can express, write it as a closed question, and run it over last month's outputs. The getting-started guide covers that afternoon.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.