03The full analysis
AI that runs the job interview: the evidence, vendors, costs, law and failure modes
In the largest randomized test of AI-led job interviews published so far, applicants interviewed by an AI voice agent were 12% more likely to receive a job offer than applicants interviewed by human recruiters, and 18% more likely to still be in the job a month later - Jabarian and Henkel, arXiv. The experiment covered 70,884 applications for customer-service jobs in the Philippines, it was pre-registered, and the hiring decisions were still made by people. When applicants were allowed to pick their interviewer, 78% chose the AI.
The same technology is pushing candidates away. In a May 2026 survey by the applicant-tracking vendor Greenhouse, 38% of US job seekers said they had walked away from a hiring process because it included an AI interview, and another 12% said they would - Greenhouse. Seven in ten said no one had told them upfront that AI would be evaluating them. And in the summer of 2025, two security researchers found that the admin console behind McDonald's AI hiring assistant accepted the password "123456", exposing a database they said covered more than 64 million applicant records - Ian Carroll and Sam Curry.
Both pictures are accurate, and the distance between them is the subject of this guide. The AI interviewer is at once one of the most rigorously tested tools in recruiting and one of the fastest ways to lose candidate trust, start a legal argument or leak a data set. Which of those outcomes a team gets depends less on the vendor it picks than on where it deploys the interviewer, how the interview is designed, what candidates are told, and who makes the decision afterwards.
This guide is written for talent leaders, recruiters and HR-technology buyers deciding whether, where and how to use an AI interviewer in late 2026. It opens with a scored comparison of the vendors, then explains from first principles why the first-round interview is the first part of hiring to be automated, walks through the field evidence and its limits, prices the options, maps the failure modes and the law as of October 2026, and ends with a pilot design and a decision framework. It builds on our State of AI in Recruiting: 2026, and it is the companion to our interview intelligence deep dive, which covers the AI notetakers that sit inside human-led interviews rather than replacing the interviewer.
Contents
- The AI interviewer scorecard
- What an AI interviewer is, and what it is not
- Why the first-round interview is the first to be automated
- The evidence: a 70,000-application field experiment
- The vendors in late 2026
- What an AI interviewer costs
- Where AI interviewers work, and where they backfire
- How AI interviewers fail
- The law in October 2026
- The candidate on the other end of the call
- How to run a pilot that survives scrutiny
- Build or buy
- Outlook: the AI-to-AI interview
- A decision framework
1. The AI interviewer scorecard
Before the detail, here is the market side by side. The scorecard ranks the AI interviewers an employer can actually buy in October 2026 on four criteria chosen from what goes wrong in practice: whether the product holds a real conversation, whether its fairness can be checked by an outsider, whether it protects candidates and their data, and whether a buyer can see what it costs. Every cell carries the evidence behind its score, and every piece of evidence comes from the vendor's own pages, a public audit or a primary filing read for this guide.
The scores measure what a buyer can verify from the outside, which is deliberately not the same as how good the product feels in a demo. A vendor with a strong product and no published audit scores lower on fairness evidence than one that publishes its audits, because a buyer cannot rely on what it cannot inspect. Two well-known names are left out on purpose: Mercor and micro1 both run AI interviews on applicants to their own expert networks rather than selling an interviewer to employers - Mercor. Where a criterion could not be verified, the cell says so and the final score is the weighted average of the criteria that could.
| # | Vendor | What It Does | Interview depth (30%) | Fairness evidence (25%) | Trust and security (25%) | Buying transparency (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | Alex | Live AI interviews over video, phone, SMS and WhatsApp | 9 - two-way interviews with follow-ups, 30+ languages | 9 - public monthly third-party audit, 90,396 candidates, all checks clear | 8 - SOC 2, ISO 27001 and ISO 42001, candidate consent flows | 5 - no public price, 33+ named ATS integrations | 8.0 |
| 2 | CodeSignal AI Interviewer | Phone or browser interviews for technical and sales roles | 8 - probes thin answers with follow-ups, 1 to 5 scored report | 6 - offers side-by-side scoring and an adverse impact study on your pilot | 7 - annual independent SOC 2 examination | 8 - public plans from $79 a month in shared credits | 7.3 |
| 3 | Classet | AI phone recruiter for frontline and healthcare roles | 6 - phone interviews 24/7 in 25+ languages, follow-ups not stated | 6 - public audit of 903,555 candidates, NYC scope only | 6 - SOC 2 Type II on the enterprise tier | 10 - $249 a month for 50 interviews, overage $5 | 6.8 |
| 4 | HackerRank Chakra | Agentic interviewer that sets a live technical problem | 9 - live canvas, observes how candidates work and reason | 3 - no public bias audit found | 8 - SOC 2, ISO 27001:2022 certified, CSA STAR | - (no readable price list; excluded from the average) | 6.8 |
| 5 | HireVue AI Interviewer | Incumbent's voice interviewer, built on Hireguide technology | 8 - two-way voice with probing follow-ups, four named languages | 6 - validation research, but its public NYC audit (2023) predates this product | 9 - SOC 2 Type 2, ISO 27001 and 27701, FedRAMP, opt-out and accommodations | 3 - no public price | 6.8 |
| 6 | Sapia.ai | Structured text chat interview with feedback to every candidate | 6 - chat only, no voice or video | 7 - independent disparate impact analysis published (2022) | 9 - ISO 42001, 27001, 27017, 27018, SOC 2 Type 2 | 5 - pay per hire, unlimited interviews, no figure published | 6.8 |
| 7 | HeyMilo | Voice, video and SMS screening for high-volume and staffing | 7 - structured voice and video, 24/7, 20+ languages | 9 - public monthly third-party audit, 22,914 candidates, all checks clear | 6 - SOC 2 badge, report type not stated | 4 - no public price, 14 named integrations | 6.7 |
| 8 | Ribbon | Video and phone interviews for hourly and frontline hiring | 8 - real-time adaptive conversations in 10 languages | 4 - claims a NYC audit, report gated behind a form | 5 - SOC 2, but its homepage says both Type I and Type II | 10 - $499 to $1,999 a month, overage $4.00 to $2.50 | 6.7 |
| 9 | Tenzo | Voice and video screening agent for staffing and corporate | 7 - phone or video interviews 24/7, schedules next rounds | 9 - public monthly third-party audit, 87,892 candidates, all checks clear | 7 - SOC 2 Type 2 and GDPR | 3 - quote only | 6.7 |
| 10 | ConverzAI | Virtual recruiter for staffing firms, voice, text and email | 6 - conversational screening, follow-up depth not stated | 6 - public audit with one "consider" flag on sex by race | 7 - SOC 2 Type II and HIPAA | 3 - no public price | 5.7 |
| 11 | Paradox (Workday) | Conversational screening and scheduling for frontline hiring | 5 - chat and text qualification screens, adaptive depth not verified | 3 - no public bias audit located | 4 - 2025 McHire exposure through a legacy admin account | 4 - sold through Workday, no public price | 4.1 |
Criteria explained. Interview depth (30%) carries the most weight because it is what a buyer is paying for: a real, adaptive conversation that gathers evidence, rather than a form with a voice. Fairness evidence (25%) rewards public, recent, independent bias audits over claims, because an AI interviewer that cannot show its selection rates is a liability under the laws in section 9. Trust and security (25%) covers certifications, candidate safeguards and incident history, because the largest documented failure in this category was a security failure, not a scoring one. Buying transparency (20%) rewards published prices and named integrations, which let a buyer benchmark a quote. Scores are 0 to 10 per criterion, and the final is the weighted average rounded to one decimal, with ties listed alphabetically.
The third-party audits in the fairness column are continuous dashboards published by Warden AI, which audits each product monthly and shows the sample size and results, for example 90,396 candidates in Alex's September 2026 audit - Warden AI. Warden sells those audits to the vendors it audits, so the dashboards are evidence of transparency more than proof of fairness, but they are the only place in this market where a buyer can see a vendor's numbers without a sales call. One disclosure of our own: AIRecruiter.co is published by HeroHunt.ai, which builds AI sourcing and outreach software. HeroHunt does not sell an AI interviewer and is not scored here.
Two patterns stand out. The products that publish prices (Classet, Ribbon, CodeSignal) are the ones built for smaller teams and frontline volume, while the enterprise products ask for a sales call, so the price transparency column is partly a proxy for target customer. And no vendor scores high on every criterion: the strongest conversationalists are not always the most transparent, and the most certified are not always the deepest interviewers. The rest of this guide explains how to weigh those trade-offs for a specific hiring problem, starting with what the category actually is.
2. What an AI interviewer is, and what it is not
An AI interviewer is software that conducts the interview itself. It asks the questions, listens to the answers, decides what to ask next, and hands the employer a structured record of the conversation: a transcript, scores against a rubric, a summary, and usually a recommendation. The defining feature is that no human interviewer sits on the employer's side of the conversation while it happens. Everything else (the channel, the voice, the avatar, the scoring model) is a design choice layered on top of that one fact.
The distinction matters because "AI interview" is now used for at least three different products, and each carries a different risk profile. A notetaker records and summarizes an interview that a human runs, which is the category our interview intelligence deep dive covers. An assessment is a test with right and wrong answers, scored the same way for everyone. An AI interviewer sits between the two: it is a conversation, so it adapts to what the candidate says, but it is also an evaluation, so it produces a judgment about a person. That combination is what makes it powerful, and it is also what puts it inside almost every hiring law written since 2020.
Five formats in use:
- One-way recorded video, scored by AI after the fact, the oldest format and the one Illinois wrote its 2020 law for - 820 ILCS 42/5
- Chat interviews, where candidates answer in text "in their own words" - Sapia.ai
- Live voice agents that hold a phone or browser conversation with "probing follow-up questions" - HireVue
- Live multichannel agents that interview over video, phone, SMS and WhatsApp - Alex
- Technical interviewers that set a real-world problem and observe how candidates work through it - HackerRank
The format determines what the software can perceive, and what it can perceive determines most of the risk.
A chat interview sees only text, so it cannot mishear an accent, but it is also the format a candidate can most easily hand to a chatbot of their own. A voice agent hears speech, which brings speech-recognition error into the evaluation and makes the experience harder for candidates with speech or hearing disabilities. A video agent sees a face, which can pull in biometric privacy law and, in the European Union, the ban on inferring emotions at work that we cover in section 9. Buyers who compare vendors on features alone tend to miss that the first and most consequential choice is the channel.
A typical live AI interview follows a predictable arc. The candidate receives an invitation by text or email, often within minutes of applying, and joins by phone or browser at a time of their choosing. The agent introduces itself as an AI, explains how the conversation will work, and moves through a guide of questions, asking follow-ups when an answer is thin or does not match the application. When the call ends, the system produces a transcript, a summary, scores against each rubric item and an overall recommendation, which land in the applicant tracking system for a recruiter to read. In the largest field experiment so far, that conversation took 10 to 20 minutes and covered up to 14 guideline topics - Jabarian and Henkel.
It also helps to be precise about what an AI interviewer usually does not do: decide.
In that experiment, the AI agent conducted the interview and a human recruiter reviewed the record and made the offer decision, and the agent disclosed at the start of each call that it was an AI - Jabarian and Henkel. Most commercial products work the same way, sending scores and transcripts into the applicant tracking system rather than rejecting anyone on their own. That boundary between evidence collection and decision is where most of the legal and ethical weight of the category sits, and it is the line a buyer should draw deliberately rather than inherit from a vendor's default settings.
Software that conducts the interview versus software that assists a human interviewer
The diagram also shows why the term "AI recruiter" causes confusion. An AI interviewer is one step inside a longer chain that can also include sourcing, outreach, scheduling and offer logistics, and products that automate several steps at once get marketed with the same words.
Our explainer on what an AI recruiter actually does breaks that chain into its steps and marks which ones act on candidates directly. For this guide, the scope is the one step where software holds the conversation and produces an evaluation, because that is the step with the most evidence, the most law and the most to go wrong.
The practical way to apply this section is to write down, before any vendor demo, three facts about the interviewer you are considering: the channel it uses, what it records (text, audio or video), and what happens to its output (a score a human reads, or a rule that advances and rejects). Those three answers decide which laws apply, which candidates it disadvantages, and how much a failure would cost, long before features or price enter the conversation.
3. Why the first-round interview is the first to be automated
Hiring is a sequence of filters, and each filter costs more per candidate than the one before it. Reading an application costs seconds, a first-round screen costs a recruiter's half hour plus scheduling, and a final loop can cost a day of senior staff time. Each stage exists to decide who deserves the next, more expensive look. The first-round screen sits exactly where the cost per candidate jumps from almost nothing to real human time, which makes it the stage where any change in the price of a conversation matters most.
That price has been under pressure from the volume side for five years. Ashby's analysis of 109 million applications and 247,000 jobs finds the average recruiter now processes 291 applications per hire, against roughly 100 in early 2021 - Ashby. No team can phone-screen 291 people for every hire, so teams ration screens, and the rationing tool is the resume: the cheapest signal in the funnel, the easiest to manufacture with a language model, and a weak predictor of performance: in the most cited revision of the selection research, years of job experience, the core content of most resumes, predicts performance with a validity of just 0.07 - Sackett et al.. The bottleneck is not that recruiters are slow. It is that the one stage which produces real evidence about a person has been too expensive to give to everyone.
An AI interviewer changes the price of that stage, not just its speed. Once a structured conversation costs a few dollars of software instead of a recruiter's half hour, the scarce resource is no longer the interview, and the logic of the funnel can flip: instead of reading 291 resumes to pick ten people to talk to, a team can talk to everyone who meets the basic requirements and read the evidence from the conversation. HireVue, which sells an AI interviewer, puts the recruiter time saved at 30 minutes per screen - HireVue, and on Ribbon's published plans the software cost works out at $2.00 to $4.99 per included interview - Ribbon pricing. Those are vendor numbers, but even heavily discounted they describe a cost gap of an order of magnitude.
The less obvious change is in the quality of the signal. A well-run AI interview is a structured interview by construction: the same questions, in the same order, scored against the same rubric for every candidate, which is the interview format the selection research ranks as the most predictive of all the procedures it reviewed, as section 11 details. Human recruiters drift. In the Philippine field experiment, the AI agent was more consistent than the average recruiter on every measure of interview structure, and more consistent than the top quarter of recruiters at following the interview guideline - Jabarian and Henkel. Consistency is not the same as accuracy, but inconsistency is a reliable way to add noise, and noise is what a screen exists to remove.
Time is the third force, and in hourly hiring it may be the strongest. An AI interviewer answers at 11 p.m. on a Sunday, when many shift workers actually apply, and the candidate who gets a same-day interview is the one who has not yet taken the job down the road. Chipotle's head of HR said the company cut time-to-hire by 75%, from 12 days to 4, with the Paradox conversational agent, in a statement published by Workday when it bought Paradox - Workday. Business hires per recruiter also reached a five-year high of five per quarter in early 2026, which means the teams adopting these tools are already running near capacity - Ashby.
There is a counterweight, and it explains most of the failures later in this guide. The first-round screen is also the first time a candidate talks to the employer, and for many candidates it is the moment they decide whether they want the job at all. Automating it saves the employer's time partly by spending the candidate's patience. Where candidates have few alternatives and value speed, that trade works for both sides. Where candidates hold the power, because their skills are scarce or their current job is good, the same trade reads as a signal that the employer does not think they are worth a person's time. The rest of this guide is, in effect, a map of where that line falls.
4. The evidence: a 70,000-application field experiment
Most claims about AI interviewers come from the companies that sell them, and they are usually before-and-after comparisons with no control group: time-to-hire fell after the tool went live, so the tool gets the credit. The single best piece of independent evidence is different in kind. Brian Jabarian and Luca Henkel ran a pre-registered natural field experiment with PSG Global Solutions, a recruitment-process-outsourcing subsidiary of Teleperformance, randomly assigning real applicants for real jobs to an AI interviewer, a human interviewer, or a choice between the two - Jabarian and Henkel. The latest version of the paper was posted to arXiv on 11 September 2026, and because its numbers moved between versions, every figure below comes from that version.
The setting was entry-level customer-service hiring in the Philippines for large US and European clients, paying roughly $280 to $435 a month. Between 7 March and 7 June 2025 the firm received 70,884 applications, of which 67,056 were eligible and randomized: 40,103 to the AI interviewer, 13,557 to human recruiters and 13,396 to a choice. The AI interview was a phone call of 10 to 20 minutes covering up to 14 guideline topics, followed by a language and analytical test of about 30 minutes, and the agent announced that it was an AI at the start of every call. Human recruiters then reviewed the record and decided who got an offer, in both arms. The study was registered in advance on the AEA RCT Registry as trial 15385 - AEA RCT Registry.
Headline results, as shares of all randomized applicants:
- Job offers: 9.73% with the AI interviewer against 8.70% with humans, a 12% relative increase
- Job starts: 6.71% against 5.65%, about 18% higher
- Still employed after one month: 5.89% against 4.98%, about 18% higher
- Retention at two to four months: 19%, 18% and 19% higher at two, three and four months
- Productivity of those hired: no meaningful difference in handling time, customer satisfaction or quality scores
The pattern matters more than any single number. If the AI had simply been more lenient, it would have produced more offers that turned into more early quits; instead the extra offers turned into extra people who started and stayed, with no drop in on-the-job performance. Among applicants who accepted an offer, AI-interviewed hires were also more likely to start (73.36% against 68.84%) and to be employed a month later (64.32% against 60.60%), which suggests the gain is not only about who got through the door. The paper reports no difference in why hires eventually left, voluntary or involuntary.
The paper's own figure shows the gap holding at every horizon it measured, from the offer through 120 days of employment.
AI versus human interviewer: offers, starts and retention
Why the AI did better: consistency, not charm
The authors' explanation is what they call controlled variance. The AI agent asked the guideline questions more completely and in a more consistent order than human recruiters did, while still adapting with follow-ups; it was more consistent than the average recruiter on every structural measure, and more consistent than the top quarter of recruiters at following the guideline - Jabarian and Henkel. The applicant's experience of the AI was not warmer: applicants rated it significantly less natural than a human call, while ratings for stress, comfort and fairness were statistically the same, and the paper's recommend-the-process score was 8.97 for the AI arm against 8.84 for humans, a difference that was not significant.
The paper's Figure 2 makes the consistency point visible. Human interviews spread across the whole range of topic coverage, while the AI's interviews cluster at near-complete coverage, with a second group of short interviews that ended early.
How consistently interviews followed the guideline
That second cluster is the study's most useful warning. In 7% of AI-led interviews the call was aborted by a technical failure of the agent, and in 5% the applicant explicitly refused to continue with an AI, according to the authors' classification of transcripts - Jabarian and Henkel. A vendor demo will never show those calls, and a buyer's pilot should count them from day one, because a system that fails one candidate in fourteen is a system that needs a human fallback built in.
The recruiters did not see it coming
Before the results were known, the firm's recruiters were surveyed about what they expected. Most predicted the AI would do the same or worse on offers, and nearly half expected worse retention; 61% expected AI-led interviews to be of lower quality. When they later scored interviews, recruiters actually rated AI-conducted interviews slightly higher on average (2.01 against 1.90 on a three-point scale), yet they relied on the AI interview score less when deciding on offers and leaned more on the separate language test - Jabarian and Henkel.
Share of recruiters at the partner firm expecting AI interviews to do worse, the same or better than human ones, surveyed before results were known
The practical reading is that the human reviewers, not the AI, were the conservative element of the system. A team adopting an AI interviewer should expect the same skepticism from its own recruiters, and should treat it as a design input rather than an obstacle: show reviewers the transcript, not only the score, and measure whether their offer decisions improve over time.
What the experiment does not show
The study's limits are as important as its results, and the authors state several of them. It covers one employer's process, one country, one job family and entry-level pay, and the authors expect the gains to hold where the attributes being screened are easy to specify (language proficiency, schedule availability, work history) and to be harder to achieve for "leadership potential, creativity, or highly specialized expertise". The choice arm also revealed negative sorting: applicants who picked the AI scored lower on the independent language and analytical tests than those who picked a human, which means a voluntary AI option changes who shows up in each channel.
Two further caveats deserve a sentence each. Brian Jabarian accepted an unpaid role as chief economist at PSG Global Solutions on 1 August 2025, after data collection ended, which the paper discloses on its title page. And headline numbers changed as the paper evolved: Chicago Booth's own coverage reported a 17% one-month retention gain from the 2025 working paper - Chicago Booth Review, the latest version reports 18%, and an earlier version's finding that applicants felt less gender discrimination with the AI does not appear in the September 2026 version. None of this undermines the core result, but it is a reminder to cite the version you read.
The rest of the evidence base
Beyond this experiment, the independent evidence thins quickly. Bias audits of AI hiring tools exist in volume, but most are commissioned by vendors and cover selection rates by sex and race rather than interview quality. Warden AI, which sells such audits, reports from more than 150 of them that audited AI systems produced an average impact ratio of 0.94 against 0.67 for the human processes they replaced, with 85% meeting the four-fifths rule - Warden AI. The same analysis found that age and disability were tested in only 5% of audits, which is a serious gap for interviewers that listen to voices and watch faces.
Two other randomized studies add useful texture. In a field experiment run on micro1's platform, recruiters who saw an AI interview report shortlisted candidates who passed a blind final human interview 46% of the time, against 29% for recruiters working from resumes alone, but 75% of invited candidates never completed the AI interview, and two of the four authors work for micro1 - Aka et al., arXiv. A second independent experiment, on one-way recorded interviews for US technology jobs, found AI scoring more predictive than recruiters' scoring while the format itself deterred applicants; section 7 covers it, because its lesson is about where to deploy. The research on bias inside language-model evaluation is more mixed than either, and section 8 sets it out alongside the evidence on speech recognition.
The way to use this evidence is to treat it as a strong prior, not a guarantee. The best study says a well-structured AI interview can match or beat human recruiters on the outcomes employers care about, for high-volume roles with clear criteria, while human reviewers keep the decision. It says nothing yet about senior roles, about candidates who can choose another employer, or about a vendor's model that changes every quarter. Those are exactly the questions a pilot has to answer for itself, and section 11 sets out how.
5. The vendors in late 2026
The vendor market splits along a line that matters more than any feature list: who owns the workflow the interview feeds into. An AI interviewer produces evidence about a candidate, and that evidence is only useful inside the system where the hiring decision gets made. That gives the companies that already own the applicant tracking system a structural advantage, and it explains why the biggest transactions in this category were acquisitions by suites rather than funding rounds for specialists.
The money tells the story more bluntly than any analyst could. Workday paid a fair value of $1.1 billion for Paradox, made up of $1.0 billion in cash and a $20 million stake it already held - Workday Form 10-Q. Mercor, which runs AI interviews on applicants to its own expert network, raised $350 million at a $10 billion valuation - Mercor. The specialist interviewer companies that sell to employers have raised a small fraction of that, in rounds of $6 million to $20 million.
US$ millions; one acquisition, one marketplace round and the largest disclosed rounds of specialist AI interviewer vendors, 2025 to 2026
Source: Workday 10-Q (Nov 2025); Mercor; Reworked; TechRSeries; Torys; Pulse 2.0
The chart is lopsided on purpose: it shows where the market believes the value of AI interviewing will be captured. It is not in the conversation itself, which voice models are commoditizing quickly, but in owning the hiring workflow (Workday) or owning the supply of vetted people (Mercor). A buyer should read that as a warning about vendor durability. The specialists below are often more capable and more transparent than the suites, but a buyer signing a multi-year contract with a company that has raised $8 million should plan for the possibility that it is acquired, pivots, or folds before the contract ends.
The suites and incumbents
Paradox, now part of Workday, is the most widely deployed conversational screening agent in frontline hiring. Its agent validates qualifications "through chat or text" and books interviews - Paradox, and when Workday announced the deal it said Paradox had powered more than 189 million AI-assisted candidate conversations - Workday. Workday completed the acquisition on 1 October 2025 and sells the product as the Workday Paradox Candidate Experience Agent - Workday completion release. It is closer to a screening and scheduling assistant than to an interviewer that probes answers, and its 2025 security incident, covered in section 8, is the main reason it ranks last in our scorecard.
HireVue, the incumbent of one-way video interviewing, moved to live conversation by acquiring the technology and team of Hireguide in March 2026, with a voice-based AI interviewer as the first milestone of the deal - HireVue. Its product page describes probing follow-up questions, per-answer scoring against role criteria, and candidate rights to opt out of the AI or request accommodations, alongside annual independent audits - HireVue AI Interviewer. The applicant tracking vendor Ashby took the same route in May 2026 by acquiring Talent Llama, an AI interviewing startup, and building the interviews into its own workflow - PR Newswire. Our analysis of the ATS market explains why the systems of record keep absorbing adjacent categories; AI interviewing is following the same path that interview notetakers followed a year earlier.
The specialists
Alex, the company formerly called Apriora, raised $20 million in 2025, a $17 million Series A led by Peak XV Partners plus a $3 million seed led by 1984 Ventures - Reworked. It interviews over video, phone, SMS and WhatsApp, says it supports 30 languages and 33 applicant tracking integrations, and claims 1,000,000 candidates interviewed - Alex. It tops our scorecard mainly because it pairs a capable product with a public, monthly third-party bias audit and an ISO 42001 certification for AI management - Alex Trust Center. The same company, under its old name, was also behind the most widely shared AI interview glitch of 2025, which section 8 describes, a reminder that the leaders in this market are still young.
Ribbon positions itself as the hiring platform for hourly and frontline work, "for everyone else" beyond white-collar and tech - Ribbon. It raised US$8.2 million led by Radical Ventures - Torys, runs video and phone interviews, and publishes its full price list, which makes it the easiest product in the market to cost out. Classet sells a phone interviewer called Joy to frontline, healthcare and skilled-trades employers, also with published prices, and claims a completion rate above 70% - Classet. Its public audit dashboard states that its AI "does not score, classify, or rank candidates" - Warden AI, while its own pricing page lists scores among the features, a contradiction a buyer should have the vendor resolve in writing.
HeyMilo interviews by voice and video and also screens by SMS and forms, around the clock in more than 20 languages - HeyMilo, with reported customers including Randstad and WilsonHCG and a reported $6 million in total funding - Pulse 2.0. Tenzo runs phone or video interviews "in any language" and schedules the next rounds - Tenzo, and its monthly audit covered 87,892 candidates in September 2026 - Warden AI. ConverzAI, built for staffing firms, raised a $16 million Series A led by Menlo Ventures - TechRSeries, and its public dashboard is the only one in this group that shows a flag, a "consider" result on intersectional bias by sex and race - Warden AI. Publishing that flag is to its credit; asking what was done about it is the buyer's job.
Sapia.ai is the outlier on format. It has run text-based chat interviews since 2018, letting candidates answer "in their own words", and says it has assessed more than 10 million candidates - Sapia.ai. It prices per hire with unlimited interviews, designed for companies that hire more than 500 people a year - Sapia.ai pricing, and it published an independent disparate impact analysis of its chat model by the statistics firm BLDS in 2022 - Sapia.ai. Text avoids the accent and face problems of voice and video, at the cost of being the easiest format for a candidate to outsource to a chatbot.
The technical interviewers
Technical hiring has its own branch of the market, because a coding screen can test work rather than talk. CodeSignal's AI Interviewer runs by phone or browser, "probes and asks follow up questions" when answers are thin, scores on a 1 to 5 scale, and the company offers side-by-side scoring against a customer's own reviewers plus an adverse impact study during a pilot - CodeSignal. HackerRank's Chakra goes further toward a work sample, spinning up a live canvas with a real-world problem and observing how candidates work, and HackerRank claims more than 500,000 interviews with an average candidate rating above 4.8 - HackerRank.
HackerRank's own pitch makes an argument worth taking seriously even from a vendor: voice-only interviewers "score how someone talks about their work, which rewards confident talkers". That is the strongest case for a work-sample design in technical roles, and it is consistent with the selection research, where job knowledge tests and work samples sit close behind structured interviews in predictive validity. Karat remains the counterexample, running technical interviews led by human interview engineers with AI-native exercises - Karat, and for senior engineering hiring that human-led model is still the norm.
The marketplaces that interview their own applicants
The largest AI interviewing operations in the world do not sell an interviewer at all. Mercor began as a hiring company that assessed candidates by analyzing interview transcripts, resumes and portfolios, and pivoted to supplying experts who train AI models - CNBC. micro1 now describes itself as a "data lab to train frontier models" - micro1, and every candidate for its own network goes through a conversational interview with its AI recruiter, Zara - micro1. These companies interview applicants at volumes no single employer reaches, and their research is useful, but two of the four authors of the main micro1 field study work for micro1, which a careful reader should keep in mind.
The practical way to use this map is to shortlist by workflow first and features second. If a suite you already run offers an interviewer, the integration and data advantages are real, and the question becomes whether its interview is as good as a specialist's. If you need depth, transparency or a specific channel, the specialists lead, and the scorecard in section 1 shows which ones publish enough for you to check their claims. Either way, ask every vendor the same questions, which section 11 lists.
6. What an AI interviewer costs
Pricing in this category follows the customer. Products built for small teams and frontline volume publish their prices and sell monthly plans with an allowance of interviews; products built for enterprises ask for a sales call and price per seat, per volume tier or per hire. That split is itself useful information: if a vendor will not publish a number, the published prices of its smaller rivals are the best benchmark a buyer has before the first quote arrives.
Four pricing models are in use. Monthly plans with interview allowances charge a fixed fee for a set number of completed interviews and a per-interview overage beyond it. Shared credits let one balance pay for assessments and AI interviews alike. Per-hire pricing charges only for candidates who are hired, with unlimited interviews. Enterprise quotes bundle interviews with integrations, single sign-on and security reviews. The published price points below are the ones a buyer can check today.
| Vendor and plan | Published price | What it includes | Unit cost (our arithmetic) |
|---|---|---|---|
| Ribbon Growth | $499 a month, billed annually | 100 interviews a month, overage $4.00 | $4.99 per included interview |
| Ribbon Business | $999 a month | 400 interviews a month, overage $3.00 | $2.50 per included interview |
| Ribbon Scale | $1,999 a month | 1,000 interviews a month, overage $2.50 | $2.00 per included interview |
| Classet Starter | $249 a month annually, $349 monthly | 50 completed interviews, overage $5.00 | $4.98 per included interview |
| Classet Scale | $999 a month | 250 interviews, overage $4.00 | $4.00 per included interview |
| CodeSignal Build / Grow | $79 / $479 a month annually | 60 / 420 annual credits for assessments and AI interviews | $15.80 / $13.69 per credit; credits per interview not stated |
| Sapia.ai | Per hire, figure not published | Unlimited interviews | Not computable |
The Ribbon figures come from its pricing page - Ribbon, the Classet figures from its own - Classet, and the CodeSignal plans from its plan page - CodeSignal. Sapia.ai describes its per-hire model without a number. The pattern across the published plans is a software cost of $2 to $5 per completed interview, falling with volume, and it is consistent with the per-minute cost of the underlying voice technology discussed in section 12.
The human baseline is easier to estimate than most teams assume. The median wage for human resources specialists, the occupation that includes recruiters, was $36.51 an hour in May 2025 - BLS. At that rate a 20-minute phone screen costs about $12 in wages alone, and a realistic 30 minutes including scheduling and notes about $18, before benefits or overhead. The authors of the micro1 field study used a fully loaded rate of $55 an hour - Aka et al., which puts a 30-minute screen at $27.50.
US dollars per completed screen: recruiter time at two cost assumptions versus published AI interviewer plan prices
On those numbers the AI screen is cheaper by a factor of three to ten, but the software price is the smallest part of the real cost, and the parts that are not on the invoice are where budgets go wrong. The first is candidate time. In the micro1 experiment, AI interviews cut recruiter workload from about 160 hours to 82, but applicants spent about 4,170 hours on the interviews, so each recruiter hour saved cost candidates roughly 53 hours - Aka et al.. The same study found that 75% of invited candidates did not complete a 30 to 40 minute AI interview. A cheaper screen that three in four candidates skip is not cheaper per hire.
The second hidden cost is the work around the interview: integrating with the applicant tracking system, writing and validating the rubric, having people review transcripts, commissioning bias audits, and handling the candidates the system fails, which was 7% of calls in the Philippine experiment. The third is drop-off among the candidates you most want, which section 7 shows is concentrated among those with options. None of these costs is fatal, but all of them scale with volume in a way the per-interview price does not.
The right unit is therefore cost per retained hire, not cost per interview. In the field experiment, AI-interviewed applicants produced 5.89 people still employed after a month for every 100 applications, against 4.98 with human interviewers - Jabarian and Henkel; for every thousand applicants, that is roughly nine more people still in the job. In a high-turnover role, where every early quit means another round of hiring and training, that retention difference is worth far more than the gap between a $3 and a $12 screen. Our recruiting ROI calculator lets a team plug in its own volumes, costs and time-to-fill to see where the break-even sits.
7. Where AI interviewers work, and where they backfire
The evidence so far points to a simple structural rule: an AI interviewer creates value when the employer's bottleneck is screening capacity and the candidate's main cost is waiting, and it destroys value when the candidate's attention is the scarce resource. Everything else (the channel, the vendor, the rubric) modulates that rule but does not overturn it. The field experiment that produced the best results was run on entry-level customer-service applicants, a population with many applicants per job, clear and checkable requirements, and every reason to prefer an interview tonight over a callback next week.
Four conditions predict a good fit. Volume has to be high relative to recruiter capacity, or the savings are trivial. The criteria have to be specifiable, because a rubric can only score what someone can write down; the authors of the Philippine study expect gains where attributes such as language proficiency, availability and work history are easy to screen, and much smaller gains for "leadership potential, creativity, or highly specialized expertise" - Jabarian and Henkel. Candidate power has to be low enough that a structured screen is not read as an insult. And the process has to tolerate a few percent of failed calls, with a human fallback ready.
| Hiring situation | Volume | Criteria easy to specify | Candidate power | Fit for an AI interviewer |
|---|---|---|---|---|
| Contact center, retail, warehouse, food service | Very high | Yes | Low | Strong, with disclosure and a human fallback |
| Staffing agencies and RPO | High | Mostly | Low to medium | Strong, with recruiter follow-up |
| Early careers and internships | High | Partly | Low | Good, if candidates can choose a human |
| Technical first screens | Medium | Yes, as a work sample | Medium to high | Good for work samples, weak for talk-only screens |
| Experienced professionals | Low to medium | Partly | High | Weak: offer it as an option, never a requirement |
| Senior leadership and executive search | Low | No | Very high | Poor |
The table's bottom half is where most of the backlash in the press comes from. Fortune found unemployed professionals refusing AI interviews outright, calling them "an added indignity" and a red flag about company culture - Fortune. A randomized experiment with about 3,300 US applicants for technology jobs found that one-way recorded interviews caused "an over 50% decrease in application continuation, including among the most qualified applicants", with the largest decline among women - Avery et al.. The same study found that a live online interview reduced continuation far less, which suggests the damage comes from the format (talking to no one) as much as from the AI.
That same paper also contains one of the strongest results in favor of AI evaluation: the commercial AI tool scored women and under-represented minorities higher than human evaluators did, and predicted applicants' employment success a year later substantially better than the recruiters. Read together, the two findings say something precise. AI scoring can be fairer and more predictive than human scoring, and a badly chosen interview format can still drive away the applicants it would have scored best. The scoring and the experience are separate design problems, and a team that solves one has not solved the other.
The practical move is to map your roles onto the table before you look at vendors. Start where the fit is strongest, which for most employers means the highest-volume, most clearly specified role, and treat any expansion upward in seniority as a separate decision with its own evidence. For experienced and senior candidates, the defensible use of AI is usually upstream of the conversation (sourcing, scheduling, notetaking) rather than in place of it.
8. How AI interviewers fail
Most writing about AI interviewer risk focuses on bias, and bias matters, but the documented failures of 2025 and 2026 are more varied and more mundane. They cluster into five types: security, reliability, perception, gaming and fairness. Each has a known mitigation, and a buyer who checks for all five will have covered most of what has actually gone wrong in production.
Security: the McHire exposure
The most consequential failure so far was not an AI failure at all. In June 2025, security researchers Ian Carroll and Sam Curry found that the administration interface of McHire, the Paradox-built hiring assistant used by most McDonald's franchisees, accepted the default credentials "123456:123456", and that a flaw in an internal API let them reach any applicant's contact details and chats; they said the data of more than 64 million applicants was exposed - Ian Carroll and Sam Curry. WIRED reported that the account had no multifactor authentication - WIRED. Paradox said the researchers accessed seven records, five containing personal information, through a legacy test account that "had not been logged into since 2019 and frankly, should have been decommissioned" - Paradox.
A week later, KrebsOnSecurity reported that a Paradox developer in Vietnam had been infected with password-stealing malware, and had at one point used the same seven-digit password for Paradox accounts tied to several Fortune 500 clients - KrebsOnSecurity. Paradox told Krebs that few of the exposed passwords were still valid and that single sign-on with multifactor authentication had been required since 2020. The lesson for buyers is structural: an AI interviewer accumulates one of the richest personal data sets in the company, recordings, transcripts and contact details of everyone who ever applied, and it is often run by a young company. Security review belongs at the top of the procurement checklist, not the bottom.
Reliability: the interviewer that loops
In May 2025, a candidate posted a clip of an AI avatar interviewer repeating the phrase "vertical bar pilates" 14 times in a row during a first-round interview for a fitness studio - 404 Media. The bot was run by Apriora, the company now called Alex, and the candidate described the experience as "really creepy" - Futurism. Viral clips exaggerate how often this happens, but the field evidence says failures are not rare: in the Philippine experiment, 7% of AI-led interviews were aborted by a technical failure of the agent. Every deployment needs a monitored failure rate, an automatic route to a human, and a promise to the candidate that a failed call will not count against them.
Perception: speech recognition, accents and disability
An interviewer that listens inherits the weaknesses of speech recognition. The best-known measurement, from 2020, found that five commercial speech recognition systems averaged a word error rate of 0.35 for Black speakers against 0.19 for white speakers - Koenecke et al., PNAS. The systems have improved since, but no 2025 or 2026 study has measured the gap inside commercial AI interviewers, which is itself a finding. In March 2025 the ACLU of Colorado filed a complaint alleging that an AI video interview used by Intuit disadvantaged a deaf Indigenous employee who had asked for human captioning as an accommodation - HR Dive; HireVue's chief executive said the complaint rested on "an inaccurate assumption about the technology used in the interview". Whatever the outcome, accommodation requests and non-voice alternatives have to be designed in, not handled as exceptions.
Gaming: the AI-versus-AI interview
Candidates have AI too. Interview Coder, a tool built to help candidates cheat in technical interviews through a window the interviewer cannot see, became the startup Cluely - TechCrunch, which went on to raise a $15 million Series A led by Andreessen Horowitz - TechCrunch. Fabric, a vendor that sells AI interviews with cheating detection, reports that 38.5% of 19,368 interviews triggered cheating flags between July 2025 and January 2026 - Fabric; flags are not proven cheating, but the trend is clear. Upstream of the interview, a Duke-led study found hidden instructions aimed at AI screeners in at least 1% of 200,000 resumes, a sevenfold rise between July 2024 and November 2025 - Duke Pratt. Our guide to hiring fraud and the verification stack covers the deepfake and proxy-candidate side of the same arms race.
Fairness: what the bias research actually says
The research on bias in language-model evaluation is real but more mixed than headlines suggest, and the mix matters for design. Across 22 models making 30,800 pairwise choices between otherwise identical CVs, female-named candidates were chosen 56.9% of the time and the first-listed candidate 63.5% of the time, while rating CVs one at a time produced only a negligible gender effect - Rozado, arXiv. A pre-registered 2026 audit found that none of 36 planned demographic contrasts survived statistical correction, but that models rewarded first-listed candidates as much as any demographic effect it measured - Vohra and Ravikiran, arXiv.
The design lesson is concrete: score each candidate alone against a written rubric, never by comparing candidates in pairs or lists, and randomize anything that has an order. The final safeguard is measurement. The US federal guideline still treats a selection rate below four-fifths of the highest group's rate as evidence of adverse impact - 29 CFR 1607.4, and the audits that check it most often skip age and disability, the classes an interviewer that hears voices is most likely to disadvantage. A pilot that measures all four, which section 11 sets out, turns fairness from a vendor claim into something the employer can see.
9. The law in October 2026
The legal map for AI interviewers moved more in the past twelve months than in the five years before, and most of the movement was delay and replacement rather than repeal. Three principles hold across almost every jurisdiction, and they are more useful than any single statute. First, the law attaches to the evaluation, not the channel: a chat, a call or a video is regulated because it produces a judgment about a person. Second, the employer stays responsible even when a vendor runs the interview, and the leading US lawsuit is testing whether the vendor is liable too. Third, almost every regime asks for the same handful of things: tell candidates, keep a human in the decision, check outcomes by protected group, and keep records.
What follows is a map, not legal advice, and every status below carries the date of the source that establishes it. For the deeper story of how liability is shifting toward vendors, our guide to AI hiring liability after Mobley v. Workday covers the case law in detail; this section focuses on what applies specifically when software holds the interview.
The European Union: one ban already in force, the heavy duties in 2027
The most immediate rule is a ban. Since 2 February 2025, the AI Act has prohibited the use of AI systems "to infer emotions of a natural person in the areas of workplace and education institutions", except for medical or safety reasons - AI Act, Article 5(1)(f). The Act defines emotion recognition as inferring emotions or intentions from biometric data such as a face or a voice, so a text chat is not caught while an interviewer that reads enthusiasm from facial expressions or confidence from vocal tone is. Law firm Lewis Silkin's reading of the Commission's guidance is that the prohibited practices in this category "extend to candidates in a recruitment cycle" - Lewis Silkin, and fines for prohibited practices reach EUR 35 million or 7% of worldwide turnover.
Everything else in the Act treats AI interviewers as high-risk, because Annex III lists systems used "to evaluate candidates". Those obligations were due in August 2026, but the AI Omnibus, in force since 27 July 2026, moved the rules for Annex III systems to 2 December 2027 - European Commission. When they apply, employers using an AI interviewer as a deployer must follow the provider's instructions, assign human oversight to trained people with real authority, inform candidates that a high-risk system is being used on them, and give a rejected candidate a "clear and meaningful" explanation of the AI's role on request. An employer that builds its own interviewer takes on the much heavier provider duties instead, a point that matters for section 12.
The United States: a patchwork, and a federal push against it
Illinois has the oldest AI interview law. Its Artificial Intelligence Video Interview Act requires an employer that uses AI to analyze recorded video interviews for Illinois-based roles to notify candidates, explain "how the artificial intelligence works", and obtain consent before the interview, and it forbids evaluating anyone who has not consented - 820 ILCS 42/5. Videos must be deleted within 30 days of a candidate's request, including copies held by vendors - 820 ILCS 42/15. The Act was written for recorded video, so live voice and chat interviews fall instead under the state's broader amendment to its Human Rights Act, in force since 1 January 2026, which bars AI use that has a discriminatory effect, bans zip codes as a proxy for protected classes, and requires notice - Crowell and Moring. The state's proposed notice rules were published in May 2026 and withdrawn in June, but the statutory duties remain in force - Kilpatrick Townsend.
New York City's Local Law 144 requires a bias audit within a year before an automated employment decision tool is used, public posting of the results, and notice to candidates - NYC DCWP. Enforcement has been weak: a December 2025 audit by the New York State Comptroller found that 75% of test calls to the city's 311 line about these tools never reached the enforcing agency, and that where the agency found one compliance issue across 32 companies, the auditors found at least 17 - DLA Piper. Weak enforcement is not the same as low risk, because the audit is also a public record that plaintiffs' lawyers can read.
Colorado's 2024 AI Act never took effect. A federal court blocked its enforcement in April 2026 in a suit brought by xAI, with the US Department of Justice intervening - McDermott Will and Schulte, and the state replaced it with SB 26-189, signed on 14 May 2026 and effective 1 January 2027, which gives candidates notice, a plain-language explanation within 30 days of an adverse decision, and meaningful human review - Colorado General Assembly. California's civil rights regulations on automated-decision systems have applied since 1 October 2025; they require employers to keep automated-decision data for four years and warn that assessments which elicit information about a disability may be an unlawful medical inquiry - California Civil Rights Department. The state's privacy regulator adds rules on automated decisionmaking technology for significant decisions, including employment, from 1 January 2027 - CPPA. Maryland requires a signed consent before an employer uses facial recognition to create a facial template during an interview - Maryland Code, Labor and Employment 3-717.
At the federal level the pressure runs the other way. Executive Order 14365 of December 2025 directed the Attorney General to set up an AI Litigation Task Force "whose sole responsibility shall be to challenge State AI laws", and it singled out Colorado's algorithmic-discrimination law by name - Federal Register. An executive order does not itself override state law, so Illinois, New York City and California rules remain enforceable unless a court or Congress says otherwise, and federal anti-discrimination law (Title VII, the ADEA and the ADA) applies regardless. In Mobley v. Workday, the court certified a nationwide age-discrimination collective in May 2025, partly denied Workday's latest motion to dismiss in June 2026, and set a class certification hearing for 9 March 2027 - CourtListener. There is no finding of discrimination yet, but the case is the reason vendors now negotiate indemnities carefully.
The United Kingdom: automated decisions allowed, with safeguards
The UK took a different route. After auditing AI recruitment tool providers and making almost 300 recommendations in 2024, including on tools that inferred gender and ethnicity from candidates' names - ICO audit findings, the regulator reported in 2026 that "many employers engaging in automated recruitment are likely relying on solely automated decisions" - ICO, Recruitment rewired.
The law has moved to accommodate that practice rather than ban it. The Data (Use and Access) Act 2025 now permits solely automated significant decisions with safeguards: candidates must be told, and must be able to contest the decision and obtain human review - ICO guidance announcement. For a UK employer, that makes the human-review path in section 10 a legal requirement for fully automated rejections, not only a courtesy.
A simplified first pass by where candidates are and what the interviewer analyzes; not legal advice
The way to apply this section is to design once, to the strictest common requirements, rather than per jurisdiction. Disclose the AI before the interview starts, not during it. Never infer emotion, personality or honesty from a face or a voice. Keep a trained person in every adverse decision and give them the transcript, not only the score. Test outcomes by sex, race, age and disability, not just the first two. Keep the records for four years, and delete recordings on request within 30 days. A process built that way satisfies most of what Brussels, Springfield, Denver, Sacramento and London ask for, and it is also the process candidates in section 10 say they want.
10. The candidate on the other end of the call
The candidate data has become consistent enough to plan around. Most job seekers have now met an AI interviewer, few of them trust it, and their objections are mostly about process rather than technology. In Greenhouse's April 2026 report, 63% of US job seekers said they had already experienced an AI interview, and 70% of those evaluated by AI said it had not been clearly disclosed before their most recent one - Greenhouse.
Trust is low even before the interview starts. Gartner's survey of 2,918 candidates found that only 26% trust AI to evaluate them fairly - Gartner, archived release.
The reasons candidates give for walking away are specific, and almost all of them are design choices. The top triggers in the Greenhouse survey were recorded video interviews scored by AI with no human present (33%), companies failing to disclose how AI would be used (27%), AI monitoring during the process (26%) and a required AI-led interview (26%), and 46% of candidates wanted the option to request a human interview instead - Greenhouse report. Yet the same survey found that 38% came away from an AI interview with a more positive view of the employer, slightly more than the 34% who came away with a more negative one. The interview is not doomed to damage the brand; a badly designed one is.
Share of US respondents, Greenhouse 2026 Candidate AI Interview Report (some figures are among those who had an AI interview)
Source: Greenhouse 2026 Candidate AI Interview Report (vendor survey, US n = 1,200)
CNBC Make It's September 2026 video shows what an AI job interview looks like from the candidate's chair and asks why so many candidates now decide it is not worth sitting through. It is worth seven minutes for anyone designing the candidate-facing side of a pilot.
How AI Is Ruining Job Interviews
Choice is the strongest design lever
The most striking candidate data comes from experiments that offered a choice. In the Philippine field experiment, 78.41% of applicants who could pick their interviewer chose the AI, 81.63% of those applying remotely and 69.30% of walk-ins - Jabarian and Henkel. When Ashby offered candidates of its customers a choice between a recruiter screen and an AI interview on their own time, 36% chose the AI - PR Newswire. The populations differ, and the release does not break the roles down, but the gap is a reminder that acceptance is a property of the candidate pool, not of the technology.
Share of candidates who picked the AI interviewer over a human recruiter, by setting
Source: Jabarian and Henkel, arXiv 2607.28222v2 (2026); Ashby via PR Newswire (May 2026)
Choice has a cost that teams rarely anticipate. In the field experiment, applicants who chose the AI scored lower on the independent language and analytical tests than those who chose a human, so a voluntary channel changes who arrives in each queue. That is not a reason to withhold choice; it is a reason to score both channels on the same rubric and to compare outcomes by channel rather than assuming the AI queue and the human queue are the same people.
Five design choices that candidates reward:
- Disclose before the interview, in the invitation, not in the first sentence of the call
- Explain what is measured and how the result is used, in plain language
- Offer a human path, including for accommodations, without penalty
- Keep it short: 10 to 20 minutes in the best field result, against 30 to 40 where three in four invitees did not finish
- Give something back, such as feedback or a faster decision
Each of those choices addresses a specific withdrawal trigger in the survey data, and none of them requires a different vendor. Sapia.ai, for example, builds personalized feedback to every candidate into its chat interview, which turns the time a candidate spends into something they keep. The length point deserves particular weight: the field experiment's 10 to 20 minute calls and the micro1 study's 30 to 40 minute interviews sit at opposite ends of the completion data, and every extra minute is paid for by the candidate. A team that implements all five choices is also, not by coincidence, most of the way to meeting the notice and human-review requirements in section 9.
11. How to run a pilot that survives scrutiny
A pilot is the only way to answer the questions the published evidence leaves open: whether an AI interviewer works for your roles, your candidates and your recruiters. It is also the record you will want to show a regulator, a plaintiff's lawyer or your own board if the tool is ever questioned. The design below borrows from the field experiment that produced the best evidence in this guide, because its structure (randomize, measure outcomes after hire, keep humans in the decision) is what made its results credible.
Start with the role, not the vendor. Pick one high-volume role whose requirements can be written down, and build the interview from a job analysis: the questions, the follow-up rules and the scoring rubric should all trace to what the job requires.
That discipline is where the value of the format comes from, and our skills-based hiring playbook covers how to turn a job analysis into criteria a rubric can score. In the most cited revision of the selection research, structured interviews have an operational validity of 0.42 against 0.19 for unstructured ones, the top rank among the selection procedures reviewed - Sackett et al., Journal of Applied Psychology. An AI interviewer that runs an unstructured conversation gives up most of that advantage.
Randomize, keep people in the decision, and measure outcomes after the hire
Randomize where you can. Assign applicants for the pilot role at random to an AI interview or a human screen, or to a choice between them, so the comparison is between similar people rather than between this quarter and last.
Keep a recruiter in every decision, and give the recruiter the transcript as well as the score; the field study found that recruiters discounted AI interview scores, and seeing the evidence is how that skepticism turns into calibration rather than distrust. CodeSignal offers to run side-by-side scoring with a customer's own reviewers during a pilot, which is the right request to make of any vendor.
Measure what matters after the hire, not just inside the funnel. The core metrics are completion and drop-off by demographic group, time from application to interview, offer rate, start rate, and retention at 30 and 90 days, plus hiring-manager ratings of the people hired.
Then test adverse impact: the federal guideline treats a selection rate below four-fifths of the highest group's rate as evidence of adverse impact, and the test should cover age and disability as well as sex and race. Track the technical failure rate separately, and route every failed call to a human within a day.
Vendor questions to settle before signing:
- Security: current SOC 2 Type II and ISO 27001 reports, multifactor authentication on every admin account, and a list of subprocessors
- Model changes: written notice before any change to the model or scoring that affects candidates
- Audit access: per-group selection data for your own candidates, not only the vendor's aggregate audit
- Data: retention limits, deletion within 30 days on request including backups, no training on your candidates without consent
- Liability: indemnities for discrimination claims and data breaches that match the risk you are taking
Those five questions map to the five failure types in section 8, and a vendor's answers are themselves a signal.
A company that cannot produce a current security report, or will not commit to notifying you before changing the model that scores your candidates, is telling you how it will behave when something goes wrong. Write the answers into the contract rather than the sales deck.
Finally, decide in advance what result would mean scale, fix or stop, and write it down before the first interview runs. A sensible bar for scaling is that the AI channel matches or beats the human channel on starts and early retention, shows no group falling below the four-fifths threshold at any stage, keeps its technical failure rate in low single digits with every failure routed to a person, and does not lose noticeably more completed applications than the human screen. Setting the thresholds first protects the pilot from the most common failure of all, which is a team that has already paid for the tool reading ambiguous results as success.
12. Build or buy
Some employers will ask whether to build their own AI interviewer, and for a few of them the answer is yes. Staffing firms, outsourcing providers and labor marketplaces run interviews at volumes where a few dollars per call adds up, and the marketplaces in section 5 already run their own. The question is worth answering from first principles, because the cost structure of the technology has changed faster than most procurement assumptions.
An AI interviewer is a voice agent with an evaluation layer on top. Building one means assembling the parts that vendors bundle, and maintaining them as the underlying models change every few months. The parts are well understood, and none of them is exotic in 2026; the difficulty lies in making them work together reliably, under load, for candidates who speak in every accent and on every kind of phone line.
Five parts make up the system. Telephony and web calling let candidates join from any phone or browser. Speech recognition has to cope with accents, background noise and interruptions. A language model runs the interview guide and decides when to follow up. Speech synthesis has to answer quickly enough to feel like a conversation rather than a walkie-talkie. And a scoring and audit layer holds the rubric, the transcripts and logs, and the connection to the applicant tracking system, which is the part that turns a conversation into evidence a recruiter can use and a regulator can inspect.
The running cost of the first four parts is now measured per minute of conversation. A ranked comparison of 20 voice and sound APIs puts the all-in cost of a voice agent at $0.13 to $0.33 per minute once orchestration, transcription, a language model, speech synthesis and telephony are stacked. That comparison is published by Founden, a company run by the same founder as this site. At those rates a 15-minute interview costs roughly $1.95 to $4.95 in raw infrastructure, by our arithmetic, which is almost exactly the range the published vendor plans in section 6 charge per completed interview.
That near-equality is the most important number in the build-or-buy decision. It means vendors are not earning their margin on the conversation itself, which is close to a commodity; they earn it on scale, on the rubric and validation work, on integrations, on the audit and security apparatus, and on absorbing legal risk. A team that builds saves little on compute and takes all of that work in-house. It also changes its legal position: under the AI Act, an organization that develops a system and puts it into service under its own name is its provider, and providers of high-risk systems carry the full set of obligations on design, documentation and conformity that deployers do not - AI Act Article 16.
The case for building is therefore about control, not cost: owning the candidate data, tuning the interview to a very specific job family, avoiding dependence on a young vendor, or integrating deeply with an in-house platform. Those are legitimate reasons for a firm whose business is interviewing at scale. For an employer whose business is something else, buying, and putting the saved engineering effort into the pilot design in section 11, is almost always the better trade.
There is also a middle path that most teams overlook: configure rather than build. Most of the vendors in section 5 let a customer write its own questions, follow-up rules and rubric, and several will score side by side with the customer's own reviewers during a pilot. That gives an employer most of the control it would get from building, the interview content and the evaluation criteria, while leaving the voice infrastructure, the security work and the audit apparatus with a vendor that spreads their cost across many customers.
13. Outlook: the AI-to-AI interview
The next eighteen months are easier to forecast than most of AI, because several of the forces are already scheduled. Colorado's replacement AI law and California's automated decisionmaking rules both take effect on 1 January 2027, the class certification hearing in Mobley v. Workday is set for 9 March 2027, and the EU's high-risk obligations for recruitment systems arrive on 2 December 2027. Any AI interviewer deployed today will spend most of its first contract under rules that are not fully in force yet, which argues for designing to them now rather than retrofitting later.
The market structure is also visible already. The three most significant deals in the category, Workday buying Paradox, HireVue buying Hireguide's technology and Ashby buying Talent Llama, were all systems of record absorbing the interview, which is the same path interview notetakers took a year earlier. Our expectation, and it is a judgment rather than a finding, is that by the end of 2027 most employers will get a basic AI interviewer as a feature of their applicant tracking system, and the independent vendors will survive by being better at something specific: a channel, a language range, a technical work sample, or transparency that the suites do not match. The broader shift toward autonomous recruiting agents, of which the interviewer is one step, is covered in our agentic AI recruiters deep guide.
The force that is hardest to forecast is the candidates' own AI. Cheating tools are funded, flag rates in AI interviews are high, and some employers have responded by bringing back at least one in-person round; Google's chief executive said in 2025 that the company was doing exactly that - Entrepreneur. The logical end point is an interview in which the employer's agent questions the candidate's agent, and neither side learns much. Long before that point, the evidentiary value of an unproctored remote conversation will fall, and the interview will shift toward what is hard to fake: live work samples, verified identity, specific personal detail probed by follow-up questions, and a human conversation at the end.
That shift changes the recruiter's job rather than ending it. If software runs the first conversation, the human work moves to designing the rubric, reviewing the evidence, handling exceptions and accommodations, and persuading the candidates the company actually wants. Those are higher-skill tasks than running the fortieth phone screen of the week, and they are the reason the entry-level end of recruiting itself is being reshaped, a dynamic our analysis of the entry-level hiring collapse traces across the wider labor market.
14. A decision framework
Everything in this guide reduces to a short sequence of questions, asked in order, where a "no" at any step is a reason to stop or redesign rather than to push on. The order matters: it puts the questions that decide value and legal exposure before the ones that decide vendor and price, which is the reverse of how most procurement processes run.
The logic of the sequence follows the structure of the argument so far. Value comes first, because an AI interviewer only pays where screening capacity is the bottleneck and the criteria can be written down. Exposure comes second, because the channel decides which laws and which candidates are affected. The candidate comes third, because a process people walk away from has no value however cheap it is. Only then do the vendor's proof and the employer's own measurement come in, and they come last because they are the steps that turn a plausible design into an accountable one.
First: high volume and clear criteria? Ask whether the role is high-volume, with requirements you can write down. If not, an AI interviewer will save little and risk much, and the honest answer is to keep the human screen.
Second: a defensible channel? Text carries the least perception risk, voice adds accent and disability risk, and video adds biometric law. Emotion inference is off the table everywhere, and in the European Union it is already banned.
Third: a path to a human? Every candidate should be able to reach a person, for accommodations, for failed calls and, ideally, as a free choice. The choice data in section 10 says many will still pick the AI when it is the faster option.
Fourth: proof from the vendor? Ask for current security reports, a recent bias audit, per-group data for your own candidates, and contract terms on model changes and deletion. A vendor that cannot provide them has told you how it will behave in a crisis.
Fifth: outcomes after the hire? Randomize if you can, and judge the tool on starts and retention, not on interviews completed. A cheap interview that produces the same hires as before has saved recruiter time, which is worth something; one that produces better hires is worth much more.
A team that can answer yes to all five is in the position of the employer in the field experiment, with a structured, disclosed, human-supervised interview whose results can be checked, and the evidence says that position can beat human screening. A team that cannot is buying the 38% walk-away rate and the headlines in section 8. The difference between those outcomes is not the vendor; it is the design. For the role-by-role data behind hiring effort, our hiring effort benchmarks by function shows where interviewer hours are spent today, which is where an AI interviewer would have to earn its place.
How we researched this guide. Every figure in this guide was read from its source in early October 2026, and each link points to the page that contains the claim. Vendor statistics are labeled as vendor claims, survey results name the organization that ran them, and the field experiment is cited at its September 2026 version because its numbers changed between drafts. The scorecard reflects only information a buyer can verify publicly; a vendor that publishes more may score differently next time. Our methodology explains how we source, label confidence and correct findings, the statistics behind our wider coverage live in the AI Recruiting Statistics Index, and every figure we track is listed with its source in our talent-market data index.
This guide reflects the AI interviewer market, evidence and law as of October 2026. Prices, product capabilities, vendor ownership and legal deadlines change quickly in this category, so verify current details before any procurement or compliance decision. Nothing in this guide is legal advice.