shipwithjev

Catalog / Agents & browsers

0553GitHub

Jev Browser: an LLM plans, Jev executes browser steps

Claude supplies a goal and text; Jev selects the browser element, action, value, and completion state while Playwright executes the step.

The public implementation and evaluation description were inspected on October 1, 2026. The builder reports 40 correct outcomes on 42 live-site tasks, with no false completion claims in the latest run. Counting and sort verification remained problematic. Free-form text comes from the calling model, and uncertain or consequential actions return a status for it to handle. The reported task results have not been independently reproduced.

Ying-Kai-Liao/jev-browserREADME ↗
# jev-browser

Browser automation where an LLM plans and **Jev** decides.

> Unofficial project, not affiliated with TypeSafe. It calls the TypeSafe System One API
> with your own API key.

The calling LLM (Claude, via MCP) says what outcome it wants, one step at a time, and hands
over any text to type. For each round of a step, code describes the page. Then one ~300 ms
[Typesafe System One](https://docs.typesafe.ai) request asks Jev several questions at once:
which element, which action, which value, and is the step done / blocked / showing an error /
about to do something irreversible. Playwright performs the action. The LLM never reads page
snapshots unless it chooses to take over.

```
Claude ── browser_do("Log in", {email, password}) ──▶ jev-browser
                                                       │  loop until done / stuck / needs confirmation
                                                       │   1. settle   (network + DOM quiet)
                                                       │   2. describe (elements, labels, state, visible text, diff, counts)
                                                       │   3. Jev      (done? error? irreversible? tool? target? value?)
                                                       │   4. act      (Playwright)
Claude ◀── { status: "done", url, actions[], done_score } ─┘
```

https://github.com/user-attachments/assets/2e688df9-4985-4854-8ebe-ba97c9d13d68

Jev only answers with probability distributions: yes/no (`noul`), pick one option (`choice`)
or a rating (`score`). It never writes text. So everything free-form comes from the caller as
candidates, and code turns disagreement or low confidence into a status the LLM can act on.

## Results

42 tasks in 16 categories on live sites (see [RESULTS.md](RESULTS.md))

Also filed under Agents & browsers

  1. 0226

    Flight search with Browser Use★

    Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓

    @gregpr07 · Agents & browsers · ~$0.004 · ~7 s

  2. 0619

    jevmem: automatic project memory for Claude Code

    After every Claude Code message, Jev decides in ~0.3 s whether it holds a decision, rule or bug worth keeping, and jevmem saves it as one line in JEVMEM.md.

    @jetwaniavinash · Agents & browsers · ~0.3 s

  3. 0618

    Playwright + Jev: tests in plain language

    Write Playwright steps like "add the most expensive item to the cart". Jev finds the right buttons and fields, fills forms and checks the result.

    Andrey Popov · Agents & browsers

  4. 0617

    toolgate: a tool-call firewall for agents

    A Claude Code and MCP hook where Jev answers seven yes/no risk questions per tool call to allow, ask or deny. Usage and benchmark numbers are in the repo.

    Jaz · Agents & browsers