Ask an experienced hiring manager whether their interviews are biased and you'll get a no, usually with something behind it. They've run hundreds of interviews and watched who worked out, and their read really is better calibrated than a first-timer's.

Bias still gets in, and it usually arrives as ordinary language: the shorthand that goes into a scorecard, or gets said out loud in a debrief. Those phrases sound like judgment and behave like preference.

What makes them decisive is that there's often nothing written down beside them. In a sample of 13,382 scorecards, the median written feedback runs to 365 words, but 26.6% of those scorecards carry five words or fewer.

Why the language does so much work

Under time pressure the brain reaches for whatever is easiest to process, and a candidate who's easy to talk to feels like a strong candidate. Fluency takes less effort to evaluate than evidence does.

The feeling gets written down as a verdict, and the verdict sounds rigorous because everyone on the panel uses the same words for it.

There's one lever in our own data that visibly moves this, so it's worth putting first. Where Metaview generated the first draft of the scorecard, 50.3% of them were submitted, against 28.6% where the interviewer had to create the scorecard themselves.

That's roughly 1.76 times the submission rate, in a capped sample of 120,000 scorecards.

It's moving from a low base. Only 31.2% of candidate interviews get a scorecard at all. Across 5.5 million interviews, about 42% of candidates who advance do so without a submitted scorecard behind them.

A missing scorecard doesn't prove anyone decided from a phrase in a debrief. It shows there was no written assessment to check the decision against, which is a different and more ordinary problem.

But it's the problem that gives the phrases their power. Where five words are all that got written, five words are all the next interviewer has to go on.

The phrases that introduce bias

None of these get said in bad faith, which is exactly why they survive. In each case the replacement is more specific than the phrase, and it's usually the moment in the interview that produced the feeling.

1. "Are they a culture fit?"

The most common cover for affinity bias. In practice "fit" resolves to "like us", which screens out the perspectives the job description asked for in the first place.

It's also the phrase least likely to have anything behind it. Values that are written down can be scored against. "Fit" gets scored against an unwritten picture of the team.

Write instead: name the value, then the evidence for it. "Pushed back on the roadmap with a senior stakeholder and brought data" is a culture-add observation somebody can argue with. "Good fit" is a preference with a job title attached.

2. "I just have a gut feeling about them."

Gut is worth listening to. After enough interviews it's sometimes picking up on something real, and telling experienced interviewers to ignore it is both patronizing and useless.

The trouble is that it turns up with its reasoning detached, and a gut feeling is reliable information about your own comfort long before it's information about their ability.

Write instead: make the gut show its working. Name the thing they said that produced the feeling. If you can't find one, you've still learned something worth knowing about the room.

3. "They didn't seem confident."

Confidence is a delivery skill and for most roles it's a minor one. Penalizing its absence lands hardest on people who were taught not to oversell, and on anyone interviewing in a second language.

Rate what the answer contained. A hesitant answer that names a real tradeoff and what they'd change beats a fluent one that names nothing.

4. "They were a bit hard to read."

Usually this describes a gap between their communication style and yours, and sometimes it describes a neurotype. It very rarely describes the work. If you couldn't read them, ask a clearer question and give them a second run at it, then score what they showed you.

5. "Would I want to grab a coffee with them?"

This is the likeability test, and people defend it hard, because it gets framed as a proxy for collaboration. Teams do have to work together, so the instinct isn't unreasonable. What the question actually measures is how closely someone resembles the people already in the room.

Write instead: ask the question you're really deciding. "Would I trust them to own this problem on the week it goes wrong?" is much closer to the job, and an interview can answer it.

6. "They seem a bit junior" (or "overqualified")

Both are labels standing in for something else: age, salary expectations, or a CV that doesn't look like the last person who held the seat. "Overqualified" is usually a private guess about whether someone will get bored, and it gets filed as an assessment of whether they can do the work.

Write instead: list the competencies the role needs and check them one at a time. Seniority is a label, so go and test the thing the label is standing in for.

7. "They remind me of [someone who worked out]."

Pattern-matching to a past hire feels like earned wisdom, and it's the fastest route to a team that keeps looking like last year's team.

It has a second problem that gets less attention: you only ever pattern-match against the people you hired. The ones you passed on who'd have been brilliant never enter the comparison, because you never found out.

Write instead: name the specific trait that made the previous person good, then test for that trait in this interview.

What actually changes the record

Watching your language works for about a week. Then a debrief overruns, the next two interviews are already stacked behind it, and the shorthand comes back. Language discipline is a personal habit, and it doesn't hold across five interviewers with different calendars.

The structural advice is familiar enough: agree on the competencies before anyone interviews, ask comparable questions, and score every candidate against the same rubric.

Our data can't tell you whether any of that produces better hires. It can tell you what happens to the written record, which is a smaller claim and the one we can stand behind.

Getting the bar written down at all is the part teams underestimate.

When you understand what you’re looking for in a candidate, it’s actually quite fuzzy in your head. Being able to articulate it is almost half the problem.”
Siadhal Magos Siadhal Magos Co-founder & CEO · Metaview

Which is why the template matters more than it sounds like it should. Metaview's Notetaker joins the interview as a visible participant and transcribes it, then writes the notes against whichever template your team picked.

The same headings come back for every candidate in the loop, and each section of those notes links back to the moment in the transcript it came from.

Choosing a notes template in Metaview, with options including recruiter screen, coding interview and role alignment
Pick the structure once and every candidate in the loop gets written up the same way.

It drafts the scorecard from that transcript, and the interviewer sets the ratings and submits it. Browser extensions autofill the objective sections of Ashby and Greenhouse scorecards. The subjective fields and the hiring recommendation stay with the person who ran the interview.

A 50.3% submission rate still means nearly half of generated scorecards never get submitted, so a generated draft removes much of the admin but someone still has to review it and press the button.

Nothing here was a controlled experiment either. Teams that switch on generated scorecards may well have been running a more disciplined process already, which means the gap isn't a measured effect of the drafting.

What a linked record buys you is narrower than "less bias", and it's still the thing this article is about. When the notes point back to the minute they came from, "they didn't seem confident" has to turn into something a colleague can go and listen to.

Want your interview notes to carry the evidence?
See how Metaview turns the conversation into structured notes.
See it live

How to put it into practice

A bias workshop won't do much here. A few habits that make the phrase impossible to file as a score will.

  • Agree on the competencies and what a strong answer sounds like in the intake call, before anyone interviews.
  • Ban verdict words in notes. "Strong", "confident" and "fit" are ratings wearing the clothes of observations. Write what the candidate said or did and let the rating follow from it.
  • Give each interviewer one or two competencies to own, so they go deep on evidence instead of trading broad impressions.
  • Score during the interview or immediately after, against the same rubric. A scorecard written three days later is a memory test.
  • Look at how recommendations spread across teams in Reports, then recalibrate the bar with regular calibration sessions.

Expect to find disagreement when you go looking, and don't read it as a broken panel. Among candidates who advanced with more than one submitted scorecard, 56.3% had at least one No or Strong No somewhere in the room.

That's a narrow population and it says nothing about panels in general. It does mean a dissent on a hire you made is a normal thing to see rather than a red flag, and what you want to know is whether the panel can point at the evidence they split on.

Metaview Reports showing scorecard recommendations broken down by department, with a positive rate for each team
Recommendation spread by team in Metaview Reports, which is usually where you find out two departments have been running different bars.

The phrases will still cross your mind, and there's nothing wrong with that. They're the compressed version of everything you've learned about interviewing, and now and then they're right. They just shouldn't be the last thing anyone can find.

See it in action

Put evidence behind every interview call.

Structured notes against your team's template, linked back to the transcript they came from.

Frequently asked

Is "culture add" just a rebrand of culture fit?

It becomes one if you never write the values down. Culture add only works when the values are specific enough to score against, so an interviewer has to point at a behavior rather than a general impression of the person. If your values are "be a good human" and "move fast", culture add will drift straight back into fit within a quarter.

What do you write on a scorecard when the interview gave you little to go on?

Write that. "Could not assess system design, we ran out of time after the project deep-dive" is a legitimate entry and it tells the next interviewer exactly what to cover. A thin interview is a finding about the interview. Filing it as a soft negative on the candidate is what turns it into bias.

How do you push back when a hiring manager says "culture fit"?

Treat it as a normal follow-up rather than a challenge. Asking what the candidate said that led them there is the sort of question that gets asked in any debrief, so it doesn't land as a compliance check. Most of the time the manager can name something specific, and when they can't, they tend to hear it themselves before you say anything.

Does structured interviewing make interviews feel robotic to candidates?

It does if you read the questions off a list and never follow up. Structure means every candidate gets comparable questions against the same competencies, and it says nothing about how conversational you are in between. Ask the same core questions every time, then follow the answer wherever it goes.

Can AI reduce interview bias?

It can make the evidence easier to find, which is a narrower claim than removing bias. Metaview's Notetaker captures what was said and organizes it, and the hiring team still sets the ratings and makes the call. No capture tool stops a panel agreeing on the wrong bar, and our own data shows the writing habit changing without showing anything about the quality of the decisions.