shipwithjev

Catalog / Research & data

0567GitHub

Jev abstract screening: compare eligibility decisions with human labels

Jev combines include/exclude choices with atomic eligibility checks to screen research abstracts against a published systematic-review dataset.

Public request code and results/full/metrics.json were inspected on October 1, 2026. On 851 Cohen ADHD title-and-abstract records, the saved results report 75.3% inclusion F1 and 83.3% recall: 70 true positives, 32 false positives, and 14 missed inclusions. These are Abstract Triage labels, not full-text eligibility or SYNERGY labels. The implementation combines Choice and Noul answers in code. This is an author-run screening evaluation, not a replacement for reviewer assessment.

PistachioAIHQ/jev-synergy-screeningREADME ↗
# Jev × Cohen — ADHD Abstract Triage

Viral life-sciences demo: screen **MEDLINE title + abstract** as **include vs exclude** with [TypeSafe Jev](https://docs.typesafe.ai/introduction) (System One), compared to **Abstract Triage** gold from [Cohen et al. 2006](https://doi.org/10.1197/jamia.M1929).

**Gold story = Abstract Triage (TIAB) only — not Article Triage, not MEDLINE PT, not SYNERGY.**

Optional mode: **Bat4RCT** r3_ship (MEDLINE PT RCT tagging) remains in the UI toggle.

## Why this is fair

| Claim | Detail |
|--|--|
| Gold | Cohen **Abstract Triage Status** (`I` = include; anything else = exclude) |
| Evidence | MEDLINE **title + abstract** fetched by PMID (same modality as gold) |
| Never | Article Triage / full-text labels · SYNERGY `label_included` |
| Topic | ADHD (N=851; 84 include ≈ 9.9%) |
| Eligibility | Oregon DERP ADHD pharmacologic review (population, listed drugs, design, outcomes, duration / pub-type excludes) |
| Encode | Jev Choice `include\|exclude` + atomic Nouls (**`cohen_adhd_r2_ship`**: codes 2–7 + monograph/imaging/formulation gates); combine in code; shown in Questions drawer |

Custom open data from the authors’ page — **not CC-BY**. Cite Cohen 2006 (see `data/COHEN_LICENSE_NOTE.md`).

## Headline metrics (include class)

Rule: **`cohen_adhd_r2_ship`**. H2H-100 has only **N=10** positives — **do not headline H2H 100%**; prefer full-851 F1.


### Film grid (stratified N=200, seed **20260917**)

| Acc | Prec | Rec | F1 | TP/FP/FN/TN | Mean lat | Est. $ |
|-----|------|-----|----|-------------|----------|--------|
| **92.0%** | **57.1%** | **80.0%** | **66.7%** | 16/12/4/168 | ~523 ms | ~$0.018 |

Film uses a stratified subsample so the grid is filmable; prevalence preserved (20/180). Metrics above are on that film set and labeled as such.

Also filed under Research & data

  1. 0620

    Jev Score: rubrics, scores and confidence

    A worked guide to Jev's Score primitive: writing a request, defining rubric levels, reading recorded probabilities and the weighted-score math.

    Jev Trader · Research & data

  2. 0614

    Bot journey classification in WebDecoy

    Sends a detected bot's last 48 request paths to Jev, which picks what it's after (prices, articles…) and how it crawls (pagination, IDs…), or unknown.

    WebDecoy · Research & data

  3. 0610

    Fake / Real: link fact-checker with Jev as judge

    Paste an article or post link. Fake / Real extracts its claims, finds outside evidence, and has Jev judge whether the evidence supports or contradicts each one.

    @DansiDanutz · Research & data

  4. 0609

    AutoRubric: rubric-based evaluation

    Combines rubric science and LLM-as-a-judge research to grade outputs with AI judges: LLMs, decision models like Jev, or both.

    @deliprao · Research & data