Blog / 49
AI Resume Screening: The Use Case That Demands Adult Supervision
Resume screening with judge verdicts: what it does well, where bias law applies, why NYC-style audit rules exist, and the design that keeps hiring human.
Most pages in this cluster open with economics; this one opens with law, because resume screening is where the decision-model pattern meets regulated, life-affecting territory. Automated employment decision tools are legally scrutinized in a growing list of jurisdictions: New York City's Local Law 144, for instance, requires bias audits and candidate notice for automated hiring tools, and it is not alone; anti-discrimination law applies to your pipeline regardless of what powers it. None of this is legal advice, and your counsel signs off before anything here touches real candidates. With the adult supervision established, here's what the pattern honestly does and doesn't offer.
What verdict-based screening does well
The volume problem is real (hundreds of applications per posting, most read for seconds) and the triage machinery applies, with the questions confined to stated, job-relevant evidence per strict question craft:
- Does the resume state experience with the required tools or domains named in the posting?
- Does stated experience meet the posted minimum (with the boundary clauses doing honest work: adjacent titles, contract work, self-employment all legislated in)?
- Are the posting's hard requirements (license, clearance, location constraint the role genuinely has) addressed?
- Is this application responsive to this role versus a mass-blast mismatch?
Deliberately absent: anything inferring protected characteristics, "culture fit", proxies thereof (name-based anything, gaps-as-negatives, school prestige), and predicted performance, which is fortune-telling with a rubric. Verdicts sort by stated evidence against stated requirements, and the honest framing is screening-in support: surfacing qualified-by-evidence candidates a tired reviewer might skip, with confidence routing per the cascade and every rejection path keeping a human decision-maker, because that's both the defensible design and, in AEDT jurisdictions, potentially the legally required one.
The audit posture (non-optional here)
Everything the production and evals pages prescribe, at maximum setting: versioned questions as your documented procedure, logged verdicts with confidence, and, specific to this domain, disparate-impact testing on your own pipeline: run outcomes across demographic slices where lawfully measurable, look for skew, and treat skew as a stop-ship, not a tuning note. Model judges inherit biases; resume text correlates with demographics in ways that make "we only judge qualifications" an empirical claim requiring evidence, not a design assertion. Candidate notice where required, human review genuinely available, and the labeling-discipline audit sample running forever. If that reads as burdensome: it is, proportionally to what's at stake, and teams unwilling to carry it should point the machinery at tickets instead, where the worst failure is a misrouted refund.
Where the same machinery is lower-stakes and immediately useful: internal mobility matching (surfacing employees whose stated skills match open roles), recruiting-ops hygiene (duplicate applications via entity resolution, completeness checks, spam and mass-blast filtering), and structured intake summaries that make human review faster without ranking anyone.
Frequently asked questions
Is AI resume screening legal?
Automated employment decision tools are regulated in several jurisdictions (bias audits, candidate notice, human-review provisions) and anti-discrimination law applies everywhere; legality depends on your design, testing, and disclosures. Involve counsel before deployment, not after.
Can it evaluate candidates fairly?
It can judge stated evidence against stated requirements consistently; fairness is an empirical property you must test for (disparate-impact analysis on your own outcomes) and maintain, not a setting you enable.
Should rejections ever be automated?
The defensible pattern keeps humans on rejection decisions, with verdicts as evidence-attached pre-work; several jurisdictions push the same direction. Screening-in support, not screening-out automation.
What's the safest first use in recruiting?
Ops hygiene: dedup, completeness, responsiveness filtering, and internal-mobility matching, value without ranking humans, while your audit posture matures per the production guide.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.