jev-scout golden-set study
Hand-labeled 25-item search-triage study with pinned rubric versions and a drift baseline; reports 88% relevance and 96% credibility with all four misses decomposed.
Hand-labeled 25-item search-triage study with pinned rubric versions and a drift baseline; reports 88% relevance and 96% credibility with all four misses decomposed.
An agent reports onboarding done; the judge verifies access and equipment claims.
everyai-com · Research & data
everyai-com
Resource
Research & data
-
-
A vendor asks for 30 minutes with no agenda; the judge picks the disposition.
everyai-com · Research & data
everyai-com
Resource
Research & data
-
-
A note-taking app silently drops edits on flaky networks; the judge grades severity.
everyai-com · Research & data
everyai-com
Resource
Research & data
-
-
A crash report with steps and logs; the judge routes it and checks reproducibility.
everyai-com · Research & data
everyai-com
Resource
Research & data
-
-