What Good Hiring Evidence Looks Like When AI Writes It
10 August 2026 · 5.6 min read · Written and published by Whizz Scribe

If AI Wrote the Hiring Note, Does It Count?
Six interviews to review. Lunch is in an hour. An AI summary lands: clean, fast, confident. But fluency hides error. A peer-reviewed study found AI summaries contained nearly 1.5% hallucinations and over 3% omissions—fabricated details and changed meanings (Nature npj Digital Medicine). Forget the abstract debate. The job for your team is to decide if that machine-written note is usable evidence. That means checking its source, job linkage, comparability, and audit trail.
Your team's priority is decision-grade evidence. A machine-written note only counts if a manager can trace each claim to something seen or heard. Connect it to the job. Compare it against the same criteria used for every other candidate. And see who signed off.
The familiar moment: a polished AI hiring summary lands in your inbox
The note says: “Strong communicator. High ownership. Great fit.”
The note looks sharp but has little substantive meaning.
The risk is that AI confidently makes claims that go beyond the available evidence. Model summaries overgeneralize what a source supports in up to 73% of cases and are nearly 5 times more likely than human summaries to make those leaps (Royal Society Open Science). In hiring, this looks like “top candidate” or “natural leader”—polished generalizations without a source.
A recruiter uploads resumes. An interview tool drafts notes. An ATS generates a score. The team is working in minutes. Then a customer (the candidate) asks for feedback. Or a founder wants the rationale. Or legal asks why one person advanced and another did not. If the note can’t show its wiring, it carries no weight.
What makes hiring evidence count when the machine wrote the first draft
The standard is old school. Even with new software, the same rules for evidence apply.
The EEOC requires proof of job-relatedness tied to observable work behaviors (EEOC). The OPM calls job analysis the foundation for any selection decision (OPM Job Analysis).
Decision-grade evidence has several key qualities:
- Role-relevant. It maps to a job requirement.
- Observable. Points to something seen, heard, or written.
- Source-backed. You can find the resume line, transcript quote, or rubric entry.
- Comparable. Every candidate gets the same structured prompts and scales, as OPM guidance on structured interviews requires (OPM Structured Interviews).
- Reviewable. A person can approve, edit, or override it. The action is logged.
This is just criterion-to-job mapping and evidence over opinion. Interrater reliability is also key. The APA defines it as agreement among raters (APA Dictionary of Psychology). If two specialists read the same transcript and give different scores, the note, despite being articulate, is useless as evidence.
The 6 fields every machine-written hiring note should include
Your team's AI notes need a template. It must force these six fields every time.
| Field | What must be recorded | Why it matters |
|---|---|---|
| Source of claim | Resume section, interview question, transcript timestamp, work sample, reference note, assessment result | Creates provenance and traceability |
| Quote vs. inference | Mark whether the line is a direct quote, paraphrase, or model inference | Stops fact and interpretation from blending |
| Job criterion mapped | The exact competency, task, or requirement the claim supports | Enforces job-requirements relevance |
| Confidence or uncertainty | High, medium, low confidence, plus what is missing or ambiguous | Makes uncertainty visible instead of buried |
| Human reviewer sign-off | Reviewer name, date, and whether they approved, edited, or overrode | Keeps a human accountable in the loop |
| Timestamp and version history | When the note was generated, edited, and finalized, plus version ID | Preserves the audit trail |
These fields turn a summary into evidence. They make compliance work faster. If a candidate, founder, or regulator asks what happened, the answers are already on the desk.
A seventh field is useful for supplier evidence when a tool generates a score. Keep the vendor’s explanation of what factors feed the score, what data was used, and what validation exists. This information belongs in the hiring record.
Weak vs decision-grade: rewrite one AI candidate summary line by line
See the difference on the page.
| Weak AI line | Decision-grade rewrite |
|---|---|
| “Great communicator and natural leader.” | “In response to the customer-escalation question, the candidate described the sequence of actions, named the stakeholders involved, and explained the resolution clearly. Source: interview Q4 transcript, 12:14–13:32. Criterion mapped: stakeholder communication. Reviewer inference: communication strength supported. Leadership not established from this answer alone.” |
| “Strong culture fit.” | “No culture-fit conclusion recorded. Evidence captured only against predefined criteria: schedule reliability, conflict handling, and documentation quality.” |
| “Has deep operations experience.” | “Resume shows 3 years supervising store opening and close procedures and weekly scheduling for a team of 12. Source: resume, experience section. Criterion mapped: shift operations and team scheduling. ‘Deep’ removed because the role benchmark for seniority is not defined in the rubric.” |
| “Top candidate.” | “Candidate received 17/20 across four structured criteria. Same questions and anchored scale were used for all interviewed candidates. Final rank remains pending human review of work sample.” |
The rewrite is longer, which is appropriate. The goal is defensibility and risk assurance for the team making the call.
Red flags that should stop a hiring decision cold
These phrases stop a hiring decision cold.
- “Great fit,” “executive presence,” “high potential,” “natural empathy,” “likely to thrive here.” These are opinion containers. They are empty unless attached to observable evidence and a role-relevant criterion.
- Any score without a factor explanation. NYC’s AEDT guidance demands that bias-audit summaries explain data sources, candidate counts, and scoring rates. A headline result without a rationale is too opaque to use (NYC DCWP FAQ).
- Comparisons without a common structure. If candidates didn't get the same questions and scoring, the output is not comparable.
- Inferences presented as facts. Research shows summary failures include fabrication, negation, and context errors (Nature npj Digital Medicine). Hiring notes do the same when they turn a vague answer into a personality judgment.
- Hidden data sources. If no one on the team knows what data the tool saw, the note is unsafe.
- Missing uncertainty. A note that never says “unclear” or “insufficient evidence” is bluffing.
A benchmark already exists. One recruiting tool provides a report card showing factors behind its score: keyword matching, years of experience, and comparison to past candidates (Dayforce Help). A report card showing the factors behind its score is an acceptable start. A grade provided without that context is not.
How to store AI-generated hiring evidence so you can reconstruct the decision later
Governance sounds heavy. In practice, this simply means saving enough information to replay the decision.
You should retain these items together in the record:
- The generated note.
- The exact version shown to reviewers.
- Source artifacts used by the tool.
- Structured scorecards.
- Human edits and overrides.
- Final disposition reason.
- Tool settings or prompt templates.
Make the record discoverable by requisition, candidate, reviewer, date, and version. If your team can’t find it in minutes, your audit trail is decorative.
Retention matters. The EU AI Act requires keeping logs for at least six months for high-risk hiring systems (EUR-Lex). The UK’s ICO says employers must record overrides of automated decisions (ICO recruitment guidance, ICO employment guidance). In NYC, employers must provide data source and retention policies on request (NYC DCWP).
Don't bury retention in default ATS settings. Decide where the record lives, who can retrieve it, and how long it stays.
A one-page checklist managers can use before they approve an AI-written hiring note
Use this checklist before an AI-written note influences any hire.
Approve only if every answer is yes:
- Does every claim have a clear source?
- Is each line marked as quote, paraphrase, or inference?
- Does every inference map to a defined job criterion?
- Did the criterion come from a job analysis or structured rubric?
- Were the same questions used for all comparable candidates?
- Does the note explain scoring with factors, not just labels?
- Is uncertainty visible where evidence is thin?
- Did a human reviewer sign off, and are overrides logged?
- Is version history preserved?
- Is the record searchable and retained?
If the answer is no, the note is draft text only. It is not hiring evidence. This standard allows a team to work fast with clean notes and make better decisions.
Sources
- Nature npj Digital Medicine (study on AI summary errors)
- Royal Society Open Science (study on summary leaps and overgeneralization)
- EEOC (job-relatedness requirements tied to observable work behaviors)
- OPM: Job Analysis
- OPM: Structured Interviews
- APA Dictionary of Psychology: Interrater reliability
- NYC DCWP AEDT FAQ
- Dayforce Help: How Candidate Grades Are Determined
- EUR-Lex: Artificial Intelligence Act (log retention requirement)
- ICO: UK GDPR guidance (recruitment and selection)
- ICO: Employment guidance (recruitment and selection)
- NYC DCWP: Automated Employment Decision Tools