Candidates the AI rates Great Fit advance at 17.2%. Poor Fits advance at 5.6%. The score and the decision move together, three to one from the top bucket to the bottom.
That looks like a model doing its job well. The data can't show that, because the recruiter sees the bucket before she decides. A number that tracks the decision that hard measures influence as much as accuracy.
Overrides run in both directions. A recruiter rejects a Great Fit when the role has shifted and the candidate's strength is now the wrong one.
She revives a Poor Fit when an unconventional background reads as a gap to the model and as range to her. Each one is a human acting on a suggestion.
That makes the override the part worth studying. What was the recruiter thinking when she went the other way, and did anyone write it down?
The score is a suggestion, the override is the decision
Inbound review answers one question at volume: which of these applications is worth a recruiter's time?
An AI agent reads every application against the role's Ideal Candidate Profile. It sorts them into ICP Fit buckets, Great through Poor, with the reasoning attached to each.
That's the part people fixate on, and it's the part that matters least to the outcome. What the model hands you is only ever a suggestion.
What happens next is a person choosing to progress the candidate or reject them. That decision shapes who you end up hiring.
The disagreements are the useful cases. A Great Fit the recruiter rejects, or a Poor Fit they revive because they saw something the application understates.
Each one is an override, and the recruiter decides.
The AI only made the first read faster. Whether the override was a good call is the part the data can't settle.
Metaview's 2026 AI and Hiring Alignment Report surveyed 505 recruiting leaders and hiring managers across North America and EMEA. It found AI already in the workflow at the teams hitting their numbers.
So how people handle the score is a more useful question than whether to have one.
What counts as an override
An override is any time the recruiter's action diverges from the AI's ICP Fit read. It covers rejecting a candidate the model rated Great or Good. It also covers progressing one it rated Okay or Poor.
Adjusting the criteria so the bucket gets re-scored counts too. Every one of those is a captured human decision, and the software reaches no verdict of its own.
The AI reads every application and sorts it. Whatever bucket a candidate lands in, it never rejects anyone. A recruiter with the right permissions makes that call. On Ashby, Greenhouse, Lever and SmartRecruiters, that decision syncs straight back to the ATS.
An override therefore points two ways. It tells you something about the candidate. It also tells you where the ICP has gone stale, and where the recruiter is carrying context the application never captured.
Read this way, the override becomes the most informative event in the review. The recruiter applies context the model may not have. The record then shows the model's score beside the recruiter's action.
How the score and the decision move together
The platform logs what the model rated each application, and what the recruiter then did with it. It doesn't log how the hire turned out, because no post-hire performance or retention data reaches the platform.
That boundary rules out the question everyone wants answered. We can't tell you which overrides were right.
What we can show is the relationship between the score and the decision.

The headline finding:
Good Fits advance at 12.1%, and Okay Fits at 7.8%. Every step down the ladder costs a candidate something.
A three-to-one spread is a strong relationship. It's also what you would expect from a queue sorted top to bottom before a human reads it. That's the problem with treating it as a verdict on the model.
Recruiters work down a ranked list. The Great Fits sit at the top and the Poor Fits sit at the bottom.
Ordering could explain some of that gap. This data can't size how much, because it holds the bucket and the decision and nothing about the order in which anyone read them.
Right now, the system's not fair. When a human recruiter decides to review applications, they appear chronologically. Once they've gotten through enough and reached out to enough, they don't look at the rest. That's unfair to the people who didn't get seen.”
His point is about the old chronological queue, and it names a risk that survives the fix. However a team orders a list, the order shapes who gets read.
Nothing in the data separates the two. If someone tells you that recruiter agreement validates a fit score, they're describing a loop.
When to trust the bucket and when to look again
That puts the weight back on the override, and on the reason behind it. The reason is the only part you control.
The cases below rest on our own judgment. The data holds fit buckets and decisions, so it can't rank situations like these. Put them in front of a panel as a starting position.
| The situation in the review | What the AI score reflects | When to trust it vs override |
|---|---|---|
| Great Fit, but you know the role just shifted | A match against the ICP built from last month's job description, before the role changed. | Override, and fix the cause. The ICP is stale, so refine it instead of correcting candidate by candidate. Future reviews then use the revised criteria. |
| Poor Fit with a non-linear background | A pattern-match against conventional track records. The model is weakest on unusual paths. | Read the whole application and trust your judgment, but log the reason. This is where a good override lives, and where a pet hunch sneaks bias back in. |
| Good Fit, strong on paper, your gut says no | The evidence in the application itself. Nothing from a conversation that hasn't happened yet. | Don't override on a hunch. Progress them to a screen, ask the question that worries you, then decide with a reason you can defend. |
| A fraud flag on an otherwise strong application | Identity and application-automation signals, separate from candidate quality. | Trust the flag enough to check it. The system surfaces the risk and the reasoning. You make the final call, and it never rejects anyone for you. |
| A whole bucket you keep rejecting | An ICP the agent is still calibrating to your definition of a good hire. | Treat the pattern as feedback. Approve the ICP refinements self-calibration suggests instead of fighting the same bucket every day. |
A good override carries a reason that would hold up if you read it back a year later.
A bad one carries a feeling that felt like a reason at the time. Bias then walks back in through the recruiter's own preferences, which is what structured review was there to prevent.
This is also where custom AI columns earn their place. Add a column for a criterion the bucket doesn't capture. The thing your gut is reacting to then becomes a reason you can defend.

None of this is a verdict on any single candidate. The review data can't show whether a candidate in any bucket became a successful hire.
The point is to make the reason for the override visible, so a year later you can tell which kind it was.
Where the override disappears
This decides which candidates reach a screen, and which strong ones get cut on a feeling. It's also where the bias you adopted AI review to remove can come back in.
The gap sits in the tooling. The review tool records the AI score and the recruiter's action, and it does that well.
It stores the bucket and it stores the action. What goes missing is the sentence explaining why the recruiter went the other way, and the reason lives in her head until she forgets it.
No system can hand you the hire outcome. The reason behind an override is free to capture though, and it's what makes a decision reviewable at all.
What to do with your own override log
Start by making the reason explicit. When a recruiter rejects a Great Fit or revives a Poor Fit, that reason belongs in the log.
Then look at where they cluster, grouped by role. In Reports, a role whose queue gets overridden again and again is telling you something about the criteria.
You won't get a scoreboard of who overrides well, and you should be suspicious of anyone selling you one. What you get is which roles draw overrides, and which fit buckets they land in. That's enough to know where to look.
The third move is the one teams skip: feed what you learn back into the ICP. A bucket you keep overriding is the criteria telling you they're out of date. Fixing the criteria beats asking the team to correct the same drift by hand every week.
One caution on all of this. A quarter of override data is a small sample for any individual reviewer, and none of it establishes who was right. Treat a pattern as a prompt for a conversation.
The score and the action taken on it already sit on one record. Record the reason for each override, then read those reasons back by role each quarter.
Bring Metaview into your hiring stack.
Live notes, structured scorecards, and ATS sync - set up in under 10 minutes.
Frequently asked
What is an AI score override in recruiting?
It's when a recruiter's decision diverges from the AI's read of an application. In Metaview, the AI sorts each candidate into an ICP Fit bucket, Great through Poor, with the reasoning attached. The recruiter then progresses or rejects the candidate. When those two point in different directions, that's an override, and the recruiter's call is the one that counts.
When should a recruiter override an AI application score?
When you carry context the model can't see and you can name the reason. The clearest cases are a stale ICP after the role has shifted, and a strong candidate whose non-linear background the model reads as a gap. Be careful with the override driven by a gut feeling and no reason behind it. That's where the bias structured review removes can come back in. A good rule: if you can't write the reason in a sentence, ask a question before you override.
Can you tell whether an AI score override improved the hire?
Not from review data alone, and be wary of anyone who says otherwise. Answering it needs post-hire performance and retention, which sits in systems the review platform never sees. What review data does show is where the score and the decision diverge, and how often. Pair that with the reason the recruiter wrote at the time, and you can judge whether an override was reasoned or reflexive. That's a different question from whether it was right.
Does Metaview auto-reject candidates?
No. Metaview never auto-rejects anyone, and there is no configuration that changes that. The AI reads each application, sorts it into an ICP Fit bucket and shows its reasoning, and stops there. Progressing or rejecting a candidate is a permission-gated action that a recruiter takes in the review queue. On Ashby, Greenhouse, Lever and SmartRecruiters, that decision then syncs back to the ATS. The product page puts it this way: Humans always decide. AI informs. You decide.
How does overriding the AI affect future application scores?
It feeds back in. Application Review self-calibrates: as recruiters progress and reject candidates, the system learns from those decisions and suggests refinements to the Ideal Candidate Profile. You review and approve any change before it takes effect. Once an approved update lands, the candidate list re-evaluates against the new ICP. So a bucket you keep overriding is a signal the criteria need fixing, and fixing them updates the read.