Artificial intelligence (AI) screening takes the first pass off your recruiters’ plate. It reads every application, or runs the screening call, and hands back a rating with a reason.

That’s a lot of ratings, every day. And your recruiters act on them, sometimes on a whole batch at once.

So someone has to check that those ratings still hold to the standard your team set. Whether a “Good fit” still means what you meant at rollout is a question with no owner until someone takes it.

That job is quality assurance (QA) for AI screening. It doesn’t need new software, since it runs on records your team already has.

And perfection isn’t the bar. Even human interviewers don’t always agree: across interviews captured on Metaview, 15.0% of interview panels split across the advance-or-reject line.¹ I would hold a screen to how often your own recruiters agree with each other, and to nothing stricter.

This guide covers what AI screening does, how teams are using it, and three reviews that keep it honest: a sample audit, a drift check, and an exception queue.

By the end, you’ll know which calls to pull each cycle, what to hold the ratings to, and which review to fix first if you can only afford one.

What is AI screening, really?

AI screening is software that does the first pass on candidates for you and hands back a rating, with the reason behind it.

It comes in two shapes:

  • Application review: The AI reads each application against the role’s criteria, sorts candidates by fit, and explains each call.
  • A screening call: An AI agent holds a screening conversation with the candidate, asks your questions, and scores each answer against a rubric your team set.

You may already run one or both. Either way, a recruiter reads the rating and decides who moves forward. This guide is about the screening call, though most of it carries over to application review.

A screening call leaves more to check than a phone screen. It leaves a rating, the reason given for it, and the call itself. A phone screen whose only record is the recruiter’s notes leaves the notes. Where consent was given and the Metaview Notetaker recorded and transcribed a phone screen, that conversation is on the record too, though the judgment behind the decision still lives in what the recruiter chose to write.

Each part of a screening call’s record is something a reviewer can check a rating against.

The QA cycle in this guide starts once candidates are taking the calls. Testing a screen before rollout is its own job, and here the first live cycle sets the baseline, because only live calls come from the people who apply.

How teams are using AI in screening today.

A screening rating is easier to trust when you can see why it was given. Two customer stories show what that looks like: one on AI application review, and one on capturing recruiters’ own calls.

Workleap: a rating a recruiter can check.

Workleap’s recruiters use Metaview Application Review, which reads each application against the role’s criteria and explains its call. Senior Recruiter Johnny Drexhage puts the time saved plainly: “It’s reduced my screening time by up to 50%.”

What won the team’s trust was the reasoning behind each rating. “That level of reasoning is something us recruiters really need,” he says.

A rating with a reason can be checked against the evidence, while a bare score can only be accepted or ignored. That reasoning is what makes QA possible.

Cockroach Labs: the evidence behind a recommendation.

At Cockroach Labs, recruiters capture their calls with the Metaview Notetaker. Recruiter Leslie Niiro describes sharing a day of call notes with a hiring manager: “You can readily output a series of notes from your calls that day and say ‘this is what we talked about, this is why I think they’re a great candidate, and here’s the evidence.’”

A screening rating should meet that standard too: a call, a reason, and the evidence behind it.

Where Metaview Screening fits.

Metaview Screening runs the screening call itself. It holds a two-way screening conversation with each candidate the team chooses to assess, asks the questions in the team’s assessment plan, and follows up where an answer needs it. Then it scores the responses against each question’s rubric. A completed call comes back with a fit rating and a documented reason grounded in the call, along with its recording and transcript, mapped question by question. Metaview captures every spoken word of the screening call, so a reviewer can check a rating against what the candidate said from the record alone.

Metaview Screening result for a sample candidate: a “Good fit” rating with its written reason, checks for AI use and misrepresentation, question-by-question evidence with a “Not covered” note, and “Reject” and “Progress” buttons for the recruiter
A completed screening call as the recruiter receives it, with the reason behind the rating and the evidence under each question. Data shown is illustrative, and the candidate shown is sample data.

For every completed call, a recruiter decides who moves forward, and Metaview does not make that decision.

Metaview Screening scores every call against the plan’s rubrics, so a weak question or a loose rubric shows up in every rating it touches.

See what a completed screening call hands your recruiters.
Metaview Screening runs the screening call and scores each response against its question’s rubric.
See it live

Why AI screening needs quality assurance.

Every rating arrives with a reason, and a recruiter acts on it. Checking a sample of those ratings is how you know the screen still means what you meant.

If results aren’t reliable, AI doesn’t reduce work. It just moves it, and recruiters end up validating, correcting, and second-guessing the output.”
Shahriar Tajbakhsh Shahriar Tajbakhsh Co-founder and Chief Technology Officer · Metaview

A QA cycle is that validation work, decided in advance: which calls get read, by whom, and where each finding goes.

What QA has to catch.

Each problem below leaves a trace in the call’s record or in the team’s own log.

What goes wrong Where it shows The review that finds it
A rating that doesn’t match its call The reason cites something the candidate never said, or skips what they did say The sample audit
A question the call can’t assess well The evidence under one of the plan’s questions comes back thin, call after call The sample audit
An edit to the assessment plan The ratings for a role shift after the edit went in The drift check
A change in who applies The ratings shift after a new job board, or after the screen moved to an earlier step The drift check
A shift with no explanation yet The ratings move and nothing in the team’s log accounts for it The drift check, then the sample audit
Recruiters who stop reading the reasons Decisions follow the rating every time, and overrides thin out The exception queue
Flags no one opens Flagged calls carry a decision and no note from the person who made it The exception queue

Why the next round can’t grade the screen.

The tempting answer key is the next round: if the candidates the screen rated highly also did well with the hiring manager, the ratings must be right. Among candidates in interviews captured on Metaview who got a scorecard in both an early and a final round, the two recommendations pointed the same way 54.4% of the time.² That measures how consistently two stages read one person, and which read was right is a question it cannot answer.

54.4%
of candidates with an early-round and a final-round scorecard, across interviews captured on Metaview, got two recommendations pointing the same way.Source: Aggregated and anonymized Metaview interview data

Later rounds also apply a bar of their own, and they see mostly the candidates a recruiter advanced. A candidate rejected on a low rating never reaches them. A later round can check the screen’s high ratings and only the overridden low ones, so the calls where a wrong rating would cost the team a candidate are the ones it never sees.

That leaves the call itself, read by a recruiter, as the one answer key that covers every rating.

How to QA AI screening, step by step.

Here’s the order to go in. Each step builds on the one before it.

Set the bar with your own recruiters.

Before you grade the screen, find out how often your recruiters agree on the same calls. Don’t borrow the panel figure for it. A split panel is a looser measure than two recruiters rating the same call, because panel interviewers often assess different things and one dissent is enough to split a panel.

So measure your own rate. Give a shared set of calls to two recruiters, have both rate every call without comparing notes, and count how often they agree, level by level. Each level’s rate is the bar the screen answers to at that level.

Decide what counts as agreement before you start. The same rating level is the strict reading. The same side of advance or reject is the looser one, and it needs the team to have written down which levels it advances. Track both.

Run a sample audit against the call itself.

Pull a fresh sample each cycle and build it in layers, so a problem in one role or with one recruiter can’t hide inside an average:

  • Every rating level: Include each level the screen gives, with the lowest over-represented.
  • Every role: Include each role the screen runs on.
  • Every recruiter: Include each recruiter who acted on a screened candidate since the last cycle.

I would give the heaviest share to the calls rated lowest. The next interviewer meets every candidate a recruiter advanced. A low-rated candidate a recruiter rejected may have no reader but the audit.

Give each call to a recruiter who did not act on that candidate, and have them rate the call first, from its recording and transcript, before they open the screen’s rating. A reviewer who reads the rating first is checking whether they can agree with it, the same pull the post on overrides of application scores describes. Shuffle the calls so they don’t arrive grouped by rating, since a run of weak calls back to back can start a reviewer judging against the queue instead of the rubric.

For every call, record two things:

  • Whether the ratings agree: Note it at the strict and the looser reading, level by level.
  • Whether the reason matches the call: A right rating resting on a wrong reason is still a finding, because the next candidate who gives that answer may land on the wrong side of it.

Where the screen falls short of your recruiters’ bar, each disagreement is a call to read. Where it clears the bar but its disagreements pile up in one role or on one question, the bar is hiding a problem there.

All of this costs recruiter hours, the thing a screen is bought to save. Set the sample size once and keep it, so each cycle compares with the last. Put the audit on a fixed schedule, so a doubt about a rating has somewhere to go.

Check the ratings for drift.

Once a cycle, for each role, set the share of calls at each rating level beside that role’s last few cycles. You’re asking whether a rating still means what it meant last cycle.

Write down in advance how big a shift has to be before it goes to the sample. For a role with few calls, count calls rather than percentages.

Keep a change log beside the ratings, one line per plan edit:

  • the date the edit went in;
  • the question it touched;
  • what changed in the question or its rubric;
  • why the team made it.

Log changes in who applies the same way. A new job board, a new region, or a screen moved to an earlier step can each shift the ratings with no plan edit behind the shift. A shift the log can’t explain goes to the sample.

An edit to a question or its rubric starts a new baseline for that role, because ratings from before and after it may not mean the same thing. Compare the ratings before the edit with their own history, start a new history after it, and read calls from each side in the next audit.

Read completion beside the ratings too. A completion figure can mislead in several ways, so check which one is in play before you act on a dip.

Where the ratings shift, the audit reads those calls next. The plan changes after those calls are read, and the log records why.

Work the exception queue.

Overrides and flagged calls go in one queue. An override is a recruiter going the other way from the rating: advancing a candidate the screen rated low, or rejecting one it rated high. Ask for a one-line reason with each override, and keep it in the same log as the plan edits.

The people making those decisions part from written recommendations elsewhere in hiring too. In interviews captured on Metaview, where an interviewer’s scorecard recommended no, the candidate still advanced 6.2% of the time.³ That can include a candidate whose other interviewers said yes, or a decision made before the scorecard was in, so it is no override rate for any screen. Read each override for what the recruiter knew before you count it against either side.

A second reader opens the call behind each override and records which of these it was:

  • a rating the call does not support;
  • context the recruiter had and the screen did not;
  • a question whose rubric no longer says what the role needs;
  • a recruiter’s decision the call does not support.

A queue can also stay empty for a whole cycle, and that is a finding in itself. It can mean a very good screen, recruiters who have stopped reading the reasons, or overrides made but never logged. The sample audit speaks to the first, and your recruiters can speak to the rest.

Flagged calls go in the same queue. On Metaview Screening, the agent flags responses that are likely AI-assisted or fraudulent, and where a team requires candidates to keep video on, it also watches for signs of reading from a second screen or a script, and for chatbot use. A flag is evidence for a person to weigh. Whoever meets a flag puts the call in the queue, and someone opens it and records what they found before a decision is made on the candidate.

In an episode of 10x Recruiting, Nolan Church and Siadhal Magos talk through the rise of AI-powered fake candidates and what teams are doing about them.

Put the three reviews on one cycle.

Each review needs an owner and a cadence:

Review When it runs Who reads What it answers
The sample audit Once a cycle, at a sample size the team sets and keeps A recruiter who did not act on those candidates, and two recruiters on a shared subset Whether ratings agree with your recruiters as often as your recruiters agree with each other, level by level, and whether each reason matches its call
The drift check Each cycle, and after any plan edit or change in who applies The talent acquisition operations lead who owns the screen Whether the ratings for a role have shifted by more than the size the team set, and what changed around them
The exception queue As overrides and flags arrive A second recruiter, with the hiring manager when a rubric is in doubt What the call showed when a recruiter went against the rating, or a flag fired

Four checks to run after each cycle.

Once a cycle closes, run these four checks before you change anything:

  • Agreement by rating level: Find the level where the screen and your recruiters disagree most, and read those calls question by question. If the disagreements share a question, that question or its rubric is the first fix. If the level holds too few calls to say, carry it forward a cycle.
  • The ratings against the change log: If they shifted by more than the size the team set and nothing logged explains it, put calls from both sides of the shift in the next audit before touching the plan.
  • Where overrides cluster: If they pile up on one recruiter, read that recruiter’s overrides against their calls first, since they may be the one reading the reasons most closely. If they cluster on one question, ask whether that criterion belongs in a later stage at all.
  • The age of open flags: If any flagged call is older than the cycle, name one owner for the queue before the next cycle opens.

Write down where each check sits after the first cycle, so the next one has something to compare against.

What to fix first when you can’t do it all.

If the first cycle can afford only part of this, I would fix the sample audit first, starting with its bottom layer.

The drift check and the queue tell the audit where to look. That bottom layer is the only place a rejected candidate’s call gets a second reader at all.

See it in action.

Audit your screening ratings against the calls themselves.

Metaview Screening returns each rating with its reason, its recording, and its transcript, so a reviewer can check the rating against the call.

Frequently asked.

How many calls should an AI screening QA sample include?

Size the sample to the hours your reviewers can give it each cycle, keep enough calls at every rating level to read that level on its own, and keep the size fixed, so one cycle’s agreement compares with the last.

Do you still need QA if the screen was tested before rollout?

A pilot answers a narrower question: whether the screen agreed with your recruiters on the calls and the plan you tested. The QA cycle asks whether it still does, on live applicants, after every change made since.

What happens when a flagged screening call turns out to be fine?

The candidate carries on like anyone else, and the log records who opened the call, what they found, and why the flag did not hold. Watch the share of flags that turn out fine from cycle to cycle, since a rise in it is a reason to read flags more closely before anyone acts on them.

Is this the bias audit New York City’s Local Law 144 requires?

The QA cycle answers a different question from that law. New York City’s Department of Consumer and Worker Protection says Local Law 144 of 2021 bars employers and employment agencies from using an automated employment decision tool unless the tool has had a bias audit within one year of its use, information about the audit is publicly available, and the required notices have gone to employees or job candidates. Whether a screening tool falls under the law in your use is a question for counsel.

What about candidates who never finish the screening call?

They leave no rating to audit, so they show up only in completion. If a role’s completion falls while its ratings hold steady, treat that as a drift finding, because the candidates who stopped finishing may differ from the ones who finished.

Sources.

¹ Aggregated and anonymized Metaview interview data: 72,753 interview panels.

² The same source, a separate query: 139,336 candidates scored in both an early and a final round.

³ The same source, a separate query: the advance rate across 201,462 scorecard recommendations of no.