Phone screens, personality tests, and AI video interviews
why the pre-interview step is breaking
By WiseWorld

The biggest gap in hiring today is the unnamed middle step between qualification and the hiring manager interview. Phone screens, personality tests, AI video interviews, questionnaires, unpaid take-homes, and AI resumes collide there. They test talk, or skills in isolation, not how people use soft skills as they do the job.
Six evidence-backed problems stack there: recruiter cost, weak prediction, AI interview dropout, resume distrust, volume strain, and prep risk.
Introduction
After the first real interview, the hiring manager says some version of: "They interviewed better than they work." Phone screens, personality tests, questionnaires, and AI video interviews filtered a lot of people. They did not measure whether those people use soft skills as they do the job.
That mismatch lives in an unnamed middle step: after the resume looks qualified, before the hiring manager meets anyone. Someone has to decide who is worth the manager's time. Most teams reach for the same tools, a phone screen, a personality test, an AI video interview, an unpaid take-home, each added to protect the manager's calendar. Extra rounds also make candidates ghost back.
Recruiters lose hours there. Good candidates drop out. Weak signals pass through. And no one can say clearly what was tested. Published research and public cases from Unilever, Amazon, and McDonald's name six problems hiding in that step.
Six hiring funnel gaps: headline statistics
Key numbers on middle-step problems, AI interview dropout, resume trust, and personality test validity.
- 6: Problems hiding in the middle of the funnel
- 38%: Candidates who quit when an AI interview is required
- 7 in 10: Job seekers now using AI on applications
- 0.19 / 1.0: How well personality tests predict performance
Key findings on the pre-interview step
Five findings on phone cost, validity, AI interview dropout, resume trust, and prep risk in the middle step.
- Six different problems land in the same spot: after the resume looks good, before the hiring manager gets involved. Most tools fix only one.
- Phone screens and personality tests are still defaults, even though they are among the weakest predictors of who will do the job well.
- 38% of candidates quit a hiring process once an AI interview is required. What saves the employer time drives candidates away.
- About seven in ten job seekers now use AI on their applications, so the resume no longer means what it used to.
- If a test uses the same questions for everyone, candidates can prepare for it. Tasks tied to the real job are much harder to game.
Six questions on the middle step
Questions the crowded middle step raises for talent acquisition teams redesigning pre-interview screening.
- Why is the step between "qualified" and the hiring manager interview so crowded?
- What does a phone screen actually cost, and what signal does it return?
- Why are candidates walking away from AI video interviews?
- Do personality tests predict job performance?
- Why did both sides stop trusting the resume?
- Which assessments still have a public answer key?
The layer nobody named
A hiring funnel diagram shows apply, screen, interview, offer. In practice, a crowded step sits between "this resume looks qualified" and "the hiring manager will meet them." Recruiters call it the phone screen, the personality test, the video interview, or just "a quick chat."
Without a name, nobody agrees what to measure. Your job post lists the skills you care about, such as teamwork or ownership. This middle step rarely checks those exact skills. In our study of European software engineer job posts, employers mostly write actions, not adjectives, yet this step still rarely tests what the post describes. Candidates prepare for what the post says. Interviewers later judge something the post never mentioned.
Unilever shows how large this step can grow. Before it rebuilt hiring, the company sorted 250,000 applications by hand to fill 800 graduate roles, which took four to six months (HireVue, 2019). Its fix was to add several tools in a row: first online games, then AI video interviews, then human interviews. People moved through faster. With each tool added, it got harder to tell candidates exactly what was being judged, and why.
Four of these six tools run in the same middle step
When each tool runs, what it does well, and where it falls short after qualification.
- Resume / AI screen (Before the middle step). Strength: Filters obvious misfits fast. Weakness: Easy to fake with AI; low trust.
- Phone screen (In the middle step). Strength: Quick human gut-check. Weakness: Eats recruiter time; weak predictor.
- Personality / AI video interview (In the middle step). Strength: Handles lots of candidates. Weakness: Candidates drop out; weak predictor.
- Skills test (In the middle step). Strength: Confirms hard skills. Weakness: Questions leak; ignores people skills.
- Soft skills in the work (In the middle step). Strength: See how they use the skills as they do the job. Weakness: Newer, not yet common.
- Hiring manager interview (After the middle step). Strength: Best read on the person. Weakness: Too slow to run on everyone.
The step nobody labels
Soft skills should be checked here, but teams pile up phone screens, personality PDFs, and AI video interviews with no shared rubric. Candidates prepare for the job post; interviewers judge something else.
- The unnamed layer: Soft skills named in posts; assessment method rarely stated (Open gap)
- The phone screen tax: High recruiter time; weak, inconsistent signal (Legacy default)
- AI video interview backlash: Scale for employers; withdrawal for candidates (Under pressure)
- Personality validity gap: Popular in hiring; weak at predicting performance (Evidence mismatch)
- Resume trust collapse: AI on both sides erodes document signal (Accelerating)
- Shared answer keys: Static tests prep-able; performance must be private (Arms race)
The phone screen tax
The phone screen is the oldest tool in this middle step, and it exists for a fair reason: a hiring manager's time is expensive, so someone has to decide who is worth a meeting. The catch is simple math. Take a mid-sized company that hires for about 60 roles a year and phone-screens roughly 30 candidates per role. That is 1,800 screens a year. At 25 minutes each, and about $45 for an hour of a recruiter's time, the phone screen alone adds up to a large bill that comes back every year.
And what do you get for it? Less than you would hope. Decades of research show that a casual, unscripted phone chat is good for breaking the ice but weak at predicting who will actually do the job well (Sackett et al., 2022). The recruiter hears a few rehearsed stories. The candidate answers "tell me about yourself" for the tenth time this month. Both spend real hours, and the result barely tells you who can do the work.
Scheduling eats the rest. In one recruiter survey, 67% said booking a single interview takes them 30 minutes to two hours, and 35% called scheduling the most time-consuming part of their job (Yello, via Interviewstream). At high volume, the phone screen becomes calendar admin long before it becomes real evaluation.
What the middle step costs in one year
Example: ~60 roles per year, ~1,800 candidates screened. Total about $79.2k ($44 per candidate). Recruiter phone screen time is the biggest line.
- Recruiter time: $46.8k
- Personality test fees: $14.4k
- AI video fees: $9.6k
- Scheduling & no-shows: $8.4k
- Total annual middle-step cost: $79.2k
For the full three-way comparison (phone vs recorded video vs self-paced job scenario), see our phone screen vs self-paced screening guide.
Candidates are walking away from AI video interviews
AI video interviews were built to fix scheduling. Tools like HireVue let a company review thousands of recorded answers without booking a single live call. Through the 2010s, more employers adopted them. Then candidates began refusing to take them.
Greenhouse's 2026 research found that 38% of candidates dropped out of a hiring process once an AI interview was required. 33% pointed to the same thing: recording answers to a camera, with an AI scoring them and no human watching. The company saves time. The candidate talks to a screen and gets a score they cannot see or question. That breaks trust, and people walk away.
Unilever is the most public example. It told The Guardian (2019) that AI video screening saved 100,000 hours of recruiter time and about $1 million worldwide in a single year. Vodafone, Intel, and Singapore Airlines used similar tools. The savings were real. So was the backlash: candidates and privacy groups objected to being scored in ways they could not see, and HireVue later dropped face-scanning from new interviews after a 2019 complaint.
The risk did not disappear. In March 2025, the ACLU filed a complaint alleging that an AI video interview at Intuit blocked a deaf employee from a promotion after her request for captions was denied. The case is still open, but it shows what happens when a tool grows faster than its fairness checks.
Candidate response to AI video interviews
Greenhouse candidate-experience research, 2026. Percentages are shares of respondents citing each outcome.
- Withdrew when AI interview was required: 38%
- One-way AI video, no human review: 33%
- Would prefer a live conversation: 61%
- Completed when the interview was optional: 72%
Most companies use video the same way: the candidate records answers to stock questions, an AI scores them, and no one explains the result. Candidates know they are performing for a machine, so many rehearse a script or simply quit.
It works differently when the interview looks like the job. Ask the candidate to handle a situation they would face in the role. Tell them what you are looking for before they start. Have a person review the recording before anyone is rejected. In that setup, there is less to rehearse and more candidates finish (IJSA, 2025; Candidate Voice Report, 2026). For a HireVue-specific breakdown of pros, cons, alternatives, and candidate sentiment, see our HireVue alternative guide.
Personality tests are popular, but weak predictors
Personality tests arrived in hiring in the 1990s and 2000s with an appealing promise: answer some questions, and we will tell you who fits. SHL, Predictive Index, Hogan, and many others built large businesses on quick quizzes and tidy reports. Buyers still like them. The evidence is not as kind.
Researchers score how well a hiring method predicts real job performance on a scale from 0 (no better than a coin flip) to 1 (perfect prediction); higher is better. In the largest review of this research, generic personality tests score about 0.19, near the bottom. A structured interview scores about 0.42, and a realistic job simulation about 0.29 (Sackett et al., 2022). In plain terms, the personality quiz is one of the weakest predictors that many teams still lean on.
Habit has outrun the evidence. A 2026 UK survey found 56% of hiring teams still use personality tests, even though only 10% of candidates think they are fair or accurate (ThriveMap). McDonald's restaurants pilot a picture-based quiz called Traitify in their hiring app; a candidate finishes it in about 90 seconds (Bersin/Paradox). Fast and easy, yes. But 90 seconds of clicking pictures tells you very little about how someone will handle a real shift.
How well each method predicts job performance
Sackett et al. (2022) meta-analysis. Taller bar = better predictor of who will do well on the job (validity × 100 for display).
- Structured interview: 42
- Job simulation: 29
- Realistic work scenario: 26
- Personality quiz: 19
- Unstructured phone screen: 19
Popular is not the same as accurate
Personality tests stay common because they are fast and familiar, not because they predict well. Structured interviews and realistic job tasks rank higher. The question is whether a stronger, job-like check can run before the manager interview.
Both sides stopped trusting the resume
Padding a resume is old. Generative AI made it effortless and universal. By Indeed's 2025–2026 research, about seven in ten job seekers now use AI to write applications or prepare for interviews. Employers answered with more AI of their own: tools that read resumes, auto-reject, and rank applicants. Candidates automate applying; employers automate filtering; and the resume stops being something either side trusts.
That reaction makes sense. A polished resume only proves someone can produce a polished resume. It does not show whether they can set priorities under pressure, disagree with a manager politely, or explain a trade-off to a colleague, the very skills most job posts ask for. So the useful evidence has to come from something a candidate cannot simply copy, paste, or generate.
Amazon learned the danger of trusting automation too much. Reuters reported in 2018 that Amazon quietly scrapped an AI tool that ranked resumes, because it had taught itself to mark down resumes that mentioned "women's" (as in "women's chess club") and graduates of women's colleges. Engineers could not confidently make it fair, so the company shut it down. Automating the resume step can scale hidden bias as fast as it saves recruiter time.
How AI changed trust in the resume
Indeed Hiring Lab (2025–2026) and Greenhouse employer surveys. Both sides now use AI, so neither fully trusts the resume alone.
- Job seekers using GenAI on applications: 70%
- Recruiters who distrust resume-only screening: 67%
- Employers adding skills tests after AI resume flood: 54%
- Candidates who say AI makes applying easier: 58%
It is an arms race, not a scandal
Candidates use AI because applying is slow and confusing. Employers use AI because they are buried in applications. The answer is to test something a candidate cannot paste into a form.
Every static test has an answer key
Once you stop trusting the resume, you lean on tests instead. But tests have a weakness: if the questions stay the same for everyone, someone will eventually work out the answers. Personality quizzes reward saying what the employer wants to hear. Skills-test questions leak onto Glassdoor and Reddit. Practiced interview stories get coached, and now drafted by ChatGPT. Research backs this up: fixed-question judgment tests can be gamed (Lievens and Dunlop, 2024), and candidates already use AI during recorded video interviews (Canagasuriam et al., 2025).
The rule of thumb is simple: if the questions never change, candidates can prepare for them. The way out is to make the task different each time and tied to the real job, so there is no fixed answer to look up. When a scenario responds to what the candidate just said, and comes from your own role, there is nothing to memorise in advance. People are also more willing to finish these tasks, because they feel like the job rather than a quiz (IJSA, 2025).
Unilever shows the trap. At global scale, every candidate played the same games and answered the same set questions. Efficient, but it created one shared target for the whole world to practice against, and coaching sites, review threads, and AI tools all teach people how to pass it.
How easy each test is to prepare for or game
Comparison index (same scale as cheat-proof assessment article). Higher bar = easier to rehearse or cheat.
- Personality quiz: prep 88, cheat 72 (indexed 0–100)
- Skills test: prep 85, cheat 68 (indexed 0–100)
- Rehearsed stories: prep 78, cheat 55 (indexed 0–100)
- Scenario quiz: prep 62, cheat 48 (indexed 0–100)
- AI interviewer: prep 58, cheat 42 (indexed 0–100)
- Soft skills in the work: prep 18, cheat 12 (indexed 0–100)
For a closer look at specific tools, see our companion piece: Can candidates cheat your soft skills assessment?.
Six gaps, one matrix
On its own, each tool has a reasonable sales pitch. Phone screens protect the manager's time. Personality tests are quick and tick a box. AI video interviews handle big volumes. Skills tests confirm technical ability. Stack them together and you get a middle step that feels busy but still sends the wrong people to the interview.
Five public employer cases — Unilever, Intuit, Amazon, McDonald's, and HireVue adopters like Vodafone — show the same loop: a tool that solved one problem quietly created a new one in trust, accuracy, or fairness.
What real employers ran into: reported wins and public challenges
Documented employer examples of AI video, personality games, and resume ranking.
- Unilever. Approach: Pymetrics games + HireVue AI video interview. Gained: Cut graduate hiring from ~4–6 months to ~4 weeks; saved 100,000 recruiter hours in one year (HireVue, 2019). Challenge: Public backlash over automated scoring; Guardian (2019) reported candidate concerns about opaque AI decisions
- Intuit. Approach: AI video interview for internal promotion. Gained: Scale across large applicant pools. Challenge: ACLU complaint (2025): deaf employee denied captioning, rejected with feedback to improve "active listening"
- Amazon. Approach: AI resume ranking tool. Gained: Automate top-of-funnel resume review. Challenge: Project scrapped by 2018: system downgraded resumes mentioning women's colleges (Reuters, 2018)
- McDonald's. Approach: Traitify picture-based personality screen in McHire. Gained: 90%+ global franchise adoption of McHire; mobile assessment in ~90 seconds (Bersin/Paradox case study). Challenge: High-volume hourly hiring still faces 130%+ industry turnover; personality screen adds another filter to game
- Vodafone, Intel, Singapore Airlines. Approach: HireVue-style AI video screening. Gained: Enterprise throughput without live scheduling. Challenge: Same category risk as Unilever: candidates cannot see how they are scored; legal scrutiny on AI hiring tools
The gap matrix compares six screening approaches on cost, validity, candidate experience, volume, and prep risk. No single column wins everything. Pick the gaps you are actually trying to close — cost, accuracy, candidate experience, volume, or cheat resistance — and check whether your current stack fixes more than one.
Six common ways to check candidates, compared side by side
Cost, validity, candidate experience, volume handling, and prep risk by method.
- Phone screen: Cost to your team High (recruiter time); Predicts performance? Weak; Candidate experience Candidates put up with it; Handles high volume? Hard at high volume; Can candidates prep for it? Rehearsed stories
- Personality quiz: Cost to your team Low per candidate; Predicts performance? Weak (0.19); Candidate experience Easy to finish; Handles high volume? Handles volume; Can candidates prep for it? Answer what they want to hear
- AI video interview: Cost to your team Medium (platform fees); Predicts performance? Medium; Candidate experience 38% drop out; Handles high volume? Handles volume; Can candidates prep for it? Rehearsed answers
- Skills test library: Cost to your team Low to medium; Predicts performance? Only hard skills; Candidate experience About 64% finish; Handles high volume? Handles volume; Can candidates prep for it? Questions leak online
- Job simulation: Cost to your team Medium (setup time); Predicts performance? Strong (0.29); Candidate experience About 83% finish; Handles high volume? Some limits at volume; Can candidates prep for it? Shared scenario libraries
- Soft skills scored in the work: Cost to your team Low per candidate; Predicts performance? Strong (job-specific); Candidate experience High when it is clear; Handles high volume? Handles volume, no live calls; Can candidates prep for it? No fixed answers to look up
How each approach evolved, and where it stalled
The phone screen predates the science of structured interviews. Personality tests spread before meta-analyses added up their weak track record. AI video interviews took off when volume was the bottleneck. Generative AI upended the resume in about three years. None of these tools are permanent fixtures — each was a fix for the problem of its day.
How each screening approach entered hiring
Timeline from phone screens through GenAI on both sides of the funnel.
- 1990s: Phone screen. Default recruiter filter before HM time.
- 2000s: Personality inventories. SHL, PI, Hogan scale to enterprise.
- 2012: AI video interviews. HireVue brings one-way video to volume hiring.
- 2018: Skills test libraries. Shared question banks go mainstream.
- 2020: Job simulations. ThriveMap, Vervoe, Harver games at scale.
- 2023+: GenAI on both sides. Candidates and employers automate the top of funnel.
Games and simulations appeared when keeping candidates engaged at volume was hard. Test libraries appeared when buyers wanted ready-made questions. AI interviewers appeared when scheduling still hurt. Every wave solved yesterday's bottleneck, not today's need: honest proof of how someone works, at scale, before the manager interview, with no answer key to look up.
What this means for talent acquisition
You do not need six new tools. You need one honest look at that middle step.
- Name the step. Write down everything that happens between "the resume looks good" and "the manager interviews them." For each part, note what it costs, how many candidates finish it, and what it actually tells you. If you cannot say what a step measures, the candidate cannot either.
- Stop paying twice for the same weak result. A phone screen and a personality test often capture the same vague impression, once by voice and once by PDF. Questionnaires and take-homes test skills in isolation. The research says seeing how they use the skills as they do the job predicts performance better than either. See our phone screen vs self-paced screening guide for the cost and validity comparison.
- Move the real test earlier, and score skills in the work. Soft skills for their own sake are meaningless. When they write, create, and hand off work from your role, there is no public list of questions to practice. You see how they use all 44 skills. That is what you wanted to know. Do not add a round. Substitute the questionnaire.
WiseWorld's take: name the middle step, then rebuild it
Six problems (cost, weak prediction, lost candidate trust, resume fraud, volume, easy cheating) pile into the same unnamed step. Name it, build the check from the real job, run it before manager time, and make it something candidates cannot look up in advance.
Further reading
Related WiseWorld research on screening methods, post-AI resumes, and the automation loop.
- Phone screen vs self-paced screening (cost for 20 candidates, completion rates)
- Pre-interview behavioral assessment guide: how they use soft skills as they do the job, before the manager interview
- Cheat-proof soft skills assessment research: validity, completion, and prep risk
- How to assess candidates after AI resume screening: post-GenAI workflow
- Hiring after AI: why candidates and recruiters keep automating each other
- HireVue alternative guide: one-way video dropout and replacement paths
Frequently asked questions
What is the biggest gap in the hiring funnel today?
The step after qualification and before the hiring manager interview is where six problems collide: weak prediction, high recruiter cost, candidate withdrawal on AI video interviews, resume distrust, questionnaires and take-homes that test skills in isolation, and tests candidates can prepare for. Most teams stack talk-based tools here instead of seeing how people use soft skills as they do the job.
Do personality tests predict job performance?
Generic personality tests show weak links to job performance, scoring about 0.19 on a 0-to-1 scale in Sackett et al. (2022) meta-analyses. Structured interviews and watching people do the work rank higher. Teams often use personality tests for speed and familiarity, not because they are the strongest predictors.
Why are candidates withdrawing from AI video interviews?
Greenhouse's 2026 research found 38% of applicants withdrew when an AI interview was required, and 33% cited one-way video scored by AI with no human review. Candidates want clarity on what is assessed and whether a human sees the recording.
How does generative AI affect resume screening?
Roughly seven in ten job seekers now use generative AI on applications and interview prep (Indeed, 2025–2026). The pile looks the same. Trust shifts from documents to how people use soft skills as they do the job, which cannot be pasted from a prompt.
Methodology
We synthesise published research and industry surveys. We do not report a new primary dataset. We prioritised sources with public methods and hiring-specific samples.
- Validity estimates: Sackett, Zhang, Berry, and Lievens (2022) meta-analysis of selection procedure validity; scores shown as relative bars (structured interview ≈ 0.42, job simulation ≈ 0.29, personality ≈ 0.19 on a 0-to-1 scale).
- Candidate experience / AI video interview withdrawal: Greenhouse Candidate Experience Report (2026); IJSA (2025) favourability ratings; Candidate Voice Report (2026) completion rates cited in our cheat-proof assessment article.
- GenAI and trust: Indeed Hiring Lab (2025–2026); Greenhouse (2025–2026); Staffing Hub AI interview report (2026).
- Assessment gaming: Lievens and Dunlop (2024); Canagasuriam et al. (2025) on AI use in AI video interviews.
- Phone screen cost model: Illustrative model for a 400-person company, ~60 roles/year, ~1,800 candidates screened in the middle step, $45/hr loaded recruiter cost, 25-minute screens, rounded for readability. Full three-way comparison in our phone screen vs self-paced screening guide. Your numbers will differ; the direction (recruiter time dominates) holds for high-volume programs.
- Funnel framing: The unnamed middle step between qualification and the hiring manager interview, where phone screens, personality tests, and AI video interviews typically stack.
- Employer case examples: Unilever/HireVue success story (2019); The Guardian (2019) on Unilever AI hiring; Reuters (2018) on Amazon resume AI; ACLU of Colorado complaint re Intuit/HireVue (2025); ThriveMap TA survey (2026); Bersin/Paradox McDonald's McHire case study on Traitify.
- Gameability index: Expert synthesis chart (same methodology as Can candidates cheat your soft skills assessment?), not a peer-reviewed metric. Useful for comparing prep risk across method families.
We interpret; we do not claim causation from cross-sectional surveys. Where vendors disagree, we cite the published number and note the limit.
More in Recruitment
Latest on the blog







