shipwithjev

Catalog / Research & data

0546Resources

I Used AI to Audit AI Bias - The Results Exposed a Systematic Pro-American Agenda in LLM…

> TL;DR: I ran 4 experiments using TypeSafe's Jev model to quantitatively measure geopolitical bias in AI recommendations. The results? 91.5% of the time, US models are placed…

> TL;DR: I ran 4 experiments using TypeSafe's Jev model to quantitatively measure geopolitical bias in AI recommendations. The results? 91.5% of the time, US models are placed first - even when Chinese models objectively outperform them on benchmarks. The bias is subtle, systematic, and hiding in pl

Also filed under Research & data

  1. 0607

    Verify: new-hire onboarding completion

    An agent reports onboarding done; the judge verifies access and equipment claims.

    everyai-com · Research & data

  2. 0606

    Triage: vague meeting request gets a disposition

    A vendor asks for 30 minutes with no agenda; the judge picks the disposition.

    everyai-com · Research & data

  3. 0605

    Triage: data-loss bug gets a severity

    A note-taking app silently drops edits on flaky networks; the judge grades severity.

    everyai-com · Research & data

  4. 0604

    Triage: crash report routing + reproducibility

    A crash report with steps and logs; the judge routes it and checks reproducibility.

    everyai-com · Research & data