REPORT
Confidently wrong
AI now touches every step of hiring. Used well it saves real time. Left unchecked, it makes critical errors that compound, step after step, into decisions you can’t trust.

THE STATE OF PLAY
AI didn’t creep into hiring. It leaped.
In two years AI has moved from a curiosity to the default way work gets done, on both candidate and employer sides of the hiring process. The time savings are real and immediate, which is exactly why the risks underneath are so often ignored.
69%
of HR teams now use AI in recruitment, up from 51% a year ago
81%
projected to be using it by 2027
39%
of candidates now use AI to apply, and climbing
Sources: SHRM, Gartner
THE CATCH
Reliability is the real barrier
Here is what the efficiency story leaves out. The single biggest thing holding teams back from AI in recruiting is not cost, and not skills, it is whether they can trust what it produces. And when an unreliable tool sits at the start of a hiring process, its errors don’t stay put. They travel all the way through it.
38%
of HR teams report significant AI impact, more than double 2023
No.2
AI is the second-biggest pressure shaping workforce decisions
57%
name AI reliability as their top barrier to using it in recruiting
Sources: Fosway Group, Talent Acquisition Realities, RecFest 2026
IT’S ALREADY HAPPENING
The problem is already here

Live poll · the webinar audience
On our recent webinar, interview notes and transcripts were the single most common answer when we asked where AI shows up most, ahead of job descriptions, CV screening and scoring. It’s also where a quiet but critical error is least likely to be caught. Here’s what that looks like up close.
The unspoken truth?
A recruiter was reviewing the AI summary notes from a batch of interviews they’d run that day when a line stopped them. The candidate, it said, brought significant experience at EY, the accountancy firm. It didn’t sound right. Nothing about EY was in the CV.
In the full transcript, what she’d actually said was UI. The AI had misheard two letters and written it up as fact, in a sentence no one ever spoke.
This one got caught. Had that summary gone to a second reviewer who wasn’t in the room, the error travels, and starts shaping a decision as if it were true. Then the question becomes: what else was misheard that nobody thought to check
AUTO GENERATED
“A strong communicator who brings significant experience in EY and a clear track record of delivery under pressure…”
Not in the CV. Not on the recording. She said “UI” not “EY”. Two letters, rewritten as a fact and passed on down the line.
109 CVs – Big red flags
Martyn Redstone, an AI and HR compliance specialist, joined our podcast to share an experiment. He ran the same batch of CVs through three different AI screening models, every day, for three weeks. The results were hard to unsee.
+2.5 places the average CV moved in the rankings, day to day
#1 → out a CV ranked top one day, gone from the top 10 the next
55% of CVs never read by the model at all before it returned a shortlist
Some drift is by design, these tools are built to vary their answers. But half the field going unread is a known limitation: when a model compares a hundred similar documents at once, it simply stops weighing them all.
WHY IT HAPPENS
Fluent is not the same as right
“An LLM is, at its heart, a very sophisticated prediction machine. Most of the time that lands well. Some of the time it is confidently, fluently wrong.”
Barny Ritchley, Chief Technology Officer
It predicts plausible words. It doesn’t know your business, and it can’t reliably tell you when it’s out of its depth.
It learns from the past, so it carries the patterns and biases of the past, including the decisions you’d rather not repeat.
And it sounds just as certain either way.
The compounding effect

Those fractions don’t just add up. They multiply.
No single stage looks like a failure on the surface but the damage compounds with every weak touchpoint.
LIVE FROM THE WEBINAR
The room couldn’t be sure
We asked the live audience whether good people could be slipping through their own process. The most common answer wasn’t yes, and it wasn’t no.
42%
couldn’t be certain their process was clean
That uncertainty is the whole problem. Unreliable AI doesn’t announce itself, so you can’t fix what you can’t see. Another 13% had already watched good people slip through, and the third who told us they don’t use AI at all still have candidates who do. The rest of this report is about removing the doubt: a simple test for any AI you rely on, and a quick check of where your own process stands.
WHERE IT HIDES
Spot it in your own data
You don’t need us to find this. Here are five tells that AI is already skewing your talent data, each one you can check against your own process this week.
Re-run a shortlist and the order changes
Same CVs, same tool, a different day, a different ranking. If it won’t sit still, it isn’t measuring the candidate
Strong applicant volume, weak fit
A job description that reads well but pulls the wrong people is writing for keywords, not for the role.
Interview notes mention things that aren’t in the CV
Spot-check a few AI summaries against the source. Misheard or invented detail travels on as fact.
A match score nobody can explain
If you can’t say why a candidate scored what they did, you can’t stand behind the decision, or answer for it.
Applications that all feel the same, arriving at scale
AI-written CVs and agent-submitted applications have quietly removed the effort that used to be the signal.
THE RELIABILITY TEST
Assist. Guard. Verify.
Three questions to run against any AI touchpoint in your hiring process, before you trust what it gives you.
1. Assist
Is AI doing the heavy lifting only where it’s genuinely strong?
Summarising, drafting, spotting patterns across more text than a person could read. Let it play to its strengths, not stray into judgment.
2. Guard
Is there a person and a guardrail where it’s weak?
Keep a human in the loop on the decisions that matter. Use AI to assist judgment, never to replace it.
3. Verify
Are you checking its output against something validated?
Don’t take its word for it. Measure against science built to predict performance, not against the AI’s own confidence.
Miss any one of these and the reliability gap opens up. Get all three right and AI becomes an asset you can stand behind.
WHAT HOLDS UP
The signals that still work
When everyone can reach for AI, the measures that survive are the ones it can’t easily solve. Here’s why, in our own data.
Aptitude
The needle barely moves
In our own data, AI produces only a small lift, on some measures, in some cohorts, and never a wide jump across the board. We still see a clear spread of high, average and low performers. The likely reason is uncertainty on the candidate’s side: there’s no easy way to tell whether the answer AI gives is even correct, and sometimes it is confidently wrong. That risk is enough to put most people off leaning on it.
Situational judgement
No answer AI can calculate
We build our SJTs so the response and scoring mechanism isn’t transparent, to a candidate or to an AI. There’s no single right answer that can be reliably worked out, so a model has nothing clean to optimize toward. It’s the design of the measure, not luck, that keeps the signal intact.
Behavioural
No template to copy
Formats like dynamic response mechanisms add complexity that’s hard to automate. And because the behaviors that matter shift from one role to the next, there’s no fixed AI template that guarantees a strong result for any given job. Success can’t be pattern-matched from the outside.
Keeping the signal clean: deter, then detect
Deter
Most won’t even try
Solid research shows that simply telling candidates they’re being monitored for AI use measurably discourages attempts. It’s easy for an assessment provider to implement, and it removes most of the problem before it starts.
Detect
The unusual stands out
We compared how candidates respond before and after AI became widely available. Responding with AI looks measurably unusual against the normal pattern, so we can flag it back to an administrator to take a closer look, rather than being quietly misled.
CHECK YOUR FUNNEL
Where is AI compounding risk for you?
Hover over the ones that sound like your process. We’ll point you to the most useful next step.
Our job descriptions are written or drafted by general AI
Your most exposed step: Job description
A weak brief is the seed for every step that follows. Our free AI Job Description Analyzer maps any job description to the competencies that actually predict success, and hands you the fixes.
TRY IT FOR FREECandidates complete assessments remotely and unsupervised
Your most exposed step: Assessment integrity
Unsupervised, remote assessment is where a minority lean on AI. Deter with an honesty commitment, then detect unusual response patterns, so you can still trust the scores.
We couldn’t fully explain how AI shaped a recent hiring decision
Your most exposed step: Explainability
Employment AI is high-risk under the EU AI Act. If you couldn’t explain a recent decision, that is the gap to close first, with science built to be measured and audited.
READ THE FULL ARTICLETHE FIX
Reliable AI at every step
Get the first step right, keep every step reliable, and the compounding runs in your favour.
1. JOB DESCRIPTION
Left to flaky AI alone:
Writes for keywords, misses what predicts success
With Saville science:
AI Job Description Analyser – maps to the competencies that matter
TRY THE AI JOB DESCRIPTION ANALYSER FREE2. ASSESSMENT DESIGN
Left to flaky AI alone:
Assess against a vague profile and you measure the wrong things
With Saville science:
Success Profiler – maps the role to the behaviours that predict success, so every score links back to what matters
BOOK A DEMO3. ASSESSMENT INTEGRITY
Left to flaky AI alone:
A minority lean on AI mid-assessment, and you can’t see who
With Saville science:
Aptitude Verification – flags unusual response patterns, so the signal stays clean
LEARN MOREHOW TO CHOOSE
Five questions to ask any AI hiring vendor
Whether you’re weighing up a new tool, your ATS, or us, these five questions separate reliable science from confident guesswork.
1. Can you show validation evidence?
Does it predict performance, or just sound right?
Ask for the data that links the tool’s output to actual job performance. Serious providers publish it. If the answer is a case study or a testimonial rather than validation evidence, treat the score as an opinion, not a measurement.
2. Can you explain a single decision?
To me, and to a regulator
Employment AI is high-risk, which means you need to be able to explain how a decision was reached. If a vendor can’t walk you through why one candidate scored above another, you inherit that gap, and the exposure that comes with it.
3. How do you handle candidates using AI?
Deter, detect, or hope?
Candidates increasingly use AI in assessments. A serious provider will tell you how they deter it up front and detect it afterwards. “Our test is AI-proof” is not an answer, nothing is, so ask what actually happens when someone tries.
4. Is this validated science, or keyword matching?
Two very different things
Many tools, most ATS matching included, rank on term frequency and keywords pulled from a job description. That’s only ever as good as the job description, and it rewards the familiar. Ask what the score is actually built on before you trust it.
5. Who checks the model for bias, and how often?
Once, or continuously?
Models learn from the past and can carry its biases. Ask who audits for adverse impact, how regularly they do it, and what happens when they find something. “We don’t see bias” usually means no one is looking.

YOUR NEXT STEP
Get it right from the first step
Try the AI Job Description Analyser
Paste in your draft job description and get a score out of 100 plus the top five fixes. Get the first step right and the rest can follow.
TRY THE AI JOB DESCRIPTION ANALYSER FREEMake AI work for you, not against you
We’d love to talk you through how we utilise reliable AI that brings benefits rather than risk, grounded in our market-leading science.
BOOK A CALLMORE RESOURCES
Podcast: From AI Hype to Audit-Ready HR · Report: HR Compliance & Audit-Ready AI