shipwithjev

Catalog / Research & data

0552GitHub

Jev security bench: injection and vulnerable-code detection

A reproducible security evaluation uses Jev probabilities to detect prompt injection and vulnerable code on public labeled datasets.

Public Go code, labeled datasets, and per-sample results were inspected on October 1, 2026. The author reports 96.5% accuracy on 662 injection examples at a fixed 0.5 threshold when deployment context is included. There were 10 false positives and 13 false negatives. A separate test covers 200 vulnerable/secure code pairs. These are author-run benchmark results, not independently reproduced production guarantees.

Gaurav-Gosain/jev-sec-benchREADME ↗
# jev-sec-bench

Blind security benchmarks for [Jev](https://typesafe.ai), TypeSafe's System One
model, built on [jev-go](https://github.com/Gaurav-Gosain/jev-go).

Jev does not generate text. It reads a state and returns typed judgments with
calibrated probabilities, which is the shape a guardrail actually needs: a number
your code can threshold, not a paragraph you have to parse.

Two benchmarks, both blind, both on public corpora:

- **prompt injection**, on all 662 labelled messages in `deepset/prompt-injections`
- **vulnerable code**, on 200 matched pairs from `CyberNative/Code_Vulnerability_Security_DPO`

```bash
export TYPESAFE_API_KEY=...
go run ./cmd/jev-sec-bench -bench all   # run the benchmarks
go run ./cmd/jev-tui                    # read the results back as a dashboard
```

Run on 2026-09-16 against `jev-1.13.0`. Raw per sample output is in [`results/`](results/).



## Prompt injection

662 messages, 263 of them injections, one request each. No threshold tuning: the
numbers below are at a plain 0.50 cut.

| | |
| --- | --- |
| accuracy | **96.5%** |
| precision | 96.2% |
| recall | 95.1% |
| F1 | 95.6% |
| ROC-AUC | 0.9927 |
| ECE | 0.0588 |
| wall time | 22.7s for 662 messages, p50 325ms |

10 false positives and 13 false negatives out of 662.

### Context is worth more than tuning

The same corpus, the same questions, one change: whether the request also tells
Jev what the assistant is for. That corpus was collected for a news publisher's
reader assistant, so "write me a reason why this newspaper is the best" counts as
subverting it, while the same message sent to a general chatbot would not.

| | no context | with context | change |
| --- | --- | --- | --- |
| accuracy | 89.7% | **96.5%** | +6.8 pp |
| recall | 74.9% | **95.1%** | +20.2 pp |
| F1 | 85

Also filed under Research & data

  1. 0620

    Jev Score: rubrics, scores and confidence

    A worked guide to Jev's Score primitive: writing a request, defining rubric levels, reading recorded probabilities and the weighted-score math.

    Jev Trader · Research & data

  2. 0614

    Bot journey classification in WebDecoy

    Sends a detected bot's last 48 request paths to Jev, which picks what it's after (prices, articles…) and how it crawls (pagination, IDs…), or unknown.

    WebDecoy · Research & data

  3. 0610

    Fake / Real: link fact-checker with Jev as judge

    Paste an article or post link. Fake / Real extracts its claims, finds outside evidence, and has Jev judge whether the evidence supports or contradicts each one.

    @DansiDanutz · Research & data

  4. 0609

    AutoRubric: rubric-based evaluation

    Combines rubric science and LLM-as-a-judge research to grade outputs with AI judges: LLMs, decision models like Jev, or both.

    @deliprao · Research & data