Yes, confident candidates tend to outscore less confident ones. In scorecards captured on Metaview, “keen” appears 2.16x as frequently per 1,000 words in a yes as in a no. That’s an uncomfortable pattern for hiring teams.
Interview scores give candidates room to win points through delivery. The written record decides whether hiring managers see the evidence beneath the impression.
I’d compare what candidates said against the same criteria, every time.
What the interview score rewards.
Candidates can change an interview score through the way they present themselves. In one study, students told to present themselves as applicants received sharply higher interview ratings than students told to answer honestly.
A correlation tells us how closely two things move together. A value closer to zero indicates a weaker relationship, while a larger positive or negative value indicates a stronger one.
A meta-analysis of 63 studies found that nonverbal cues stayed linked to interview ratings across different levels of structure, formats, and durations. Interviewers may read appearance, eye contact, and head movement as confidence, although those cues don’t measure confidence itself.
Recent work on authenticity points in the same direction. Researchers found that the way candidates sounded and behaved carried the strongest relationship with interview performance, while only verbal authenticity cues were related to ratings of their work.
Good structured interviews keep the competencies and rating method steady. They’re worth doing. But structure alone doesn’t stop delivery from counting.
The words that ride along with a yes.
Positive verdicts use positive language. The useful part is seeing which words interviewers reach for once they favor a candidate.
In submitted scorecards captured on Metaview, “adaptable” appears 1.96x as frequently per 1,000 words in a yes as in a no. “Thoughtful” appears 1.83x as frequently, and “personable” appears 1.80x as frequently.
These words show how the verdict gets written. They describe the candidate without showing what led the interviewer there.
“Adaptable” could refer to someone who changed a plan after new evidence emerged. It could also refer to someone who handled the conversation smoothly. The adjective supports both readings by itself.
Teams examining interviewer bias should ask what sits behind each trait. If the scorecard includes the candidate’s example, another reviewer can check the label against the response.
That’s the dry problem with shorthand: the adjective arrives finished. The work that earned it may never reach the page.
Nerves and confidence aren’t the same variable.
An anxious candidate can sound less assured. Anxiety and confidence still describe different things.
A 2018 meta-analysis put the correlation between self-reported anxiety and interview performance at -.19. That’s a small relationship by general standards and a medium one within this area of study.
Researchers have also tested what observers see. In a study of 823 observers, people watched an actor deliver scripted responses. They gave lower interview ratings when the actor showed highly anxious nonverbal behavior.
Self-presentation can sit beside anxiety too. Among applicants for a campus job, self-reported anxiety correlated .30 with deceptive self-presentation and -.18 with honest self-promotion.
For an interviewer, visible composure is ambiguous. It may reflect preparation or familiarity with interviews. It may also show how that candidate responded to that particular setting.
A clear candidate experience tells candidates what to expect. Interviewers still need to judge the answer against the role without giving composure an unnamed share of the score.
What the job-side evidence does and doesn’t support.
Recent studies don’t carry the interview advantage into supervisors’ ratings in the same way.
Researchers studying applicants for residence assistant positions found that interview anxiety had a near-zero relationship with ratings from supervisors and supervisees. Another study of applicants in temporary roles found a near-zero relationship between interview anxiety and supervisors’ ratings.
Research on authenticity drew a narrower line. The way candidates sounded and behaved was especially relevant to interview ratings, while the verbal cues alone were related to ratings of their work.
Confident delivery can also prompt suspicion that someone is faking. A review found deceptive self-presentation effectively unrelated to interview ratings on average. Another review found mixed effects on interview scores and said detection was hardly possible.
There’s a reasonable objection. The same review found that deceptive self-presentation didn’t consistently weaken the relationship between an interview and an outside criterion. A hiring team can’t assume every attempt to present well is empty theater.
There’s no concrete evidence that candidates who present themselves confidently are less competent than candidates who don’t. The interview score alone can’t settle that comparison.
The record most interviews leave behind.
Metaview’s artificial intelligence (AI) generated scorecards fill 3x more fields than manually written ones in an analysis limited to 80,000 scorecards.
Across a separate sample limited to 120,000 created scorecards, 28.6% of manual scorecards get submitted against 50.3% of AI-generated ones. The other cells describe different sets of captured records.
Across the interviews captured on Metaview, only 31.2% received a scorecard at all. Across 13,382 written feedback records, 26.6% ran five words or fewer.
Among a subset of 13,373 created scorecards, 32.1% had every field filled, 25.6% were completely empty, and 27.5% contained one field or fewer. Across those created scorecards, about 7 of 15 fields were filled on average.
Interviewers also disagree as candidates move through a process. Early-round and final-round scorecards agreed 54.4% of the time across 139,336 records. Agreement is separate from accuracy, and roughly half can’t be treated as a coin flip without an assumed chance baseline.
Among candidates who advanced through a multi-interviewer process, 56.3% had at least one dissenting scorecard in the captured record. The disagreement is visible. Its cause depends on what the interviewers wrote down.
The Metaview Notetaker joins an interview as a visible participant. It records and transcribes the conversation, captures every spoken word, and turns the call into structured notes. It doesn’t evaluate whether the interview was good.
Notes preserve material. Hiring teams still have to connect it to their criteria. A useful scorecard leaves an inspectable record before the panel reaches a hire or no-hire decision.
What I’d rather compare.
I’d stop comparing how convincing two candidates felt and compare the evidence they gave against the same standard.
Define a strong answer before the interview. Each question needs its own rubric, and interviewers need to know what evidence belongs under each rating. A broad competency label gives a general impression too much work to do.
Teams can build interview rubrics at the question level, then ask questions that surface comparable evidence. A follow-up can retrieve a missing part of the example.
If an interviewer writes “pragmatic,” the scorecard should also record the constraint the candidate faced and the result they reported. Another reviewer then has something concrete for assessing candidate responses.
Interviewers should set their ratings before the interview debrief. Each person completes the relevant fields and links the rating to recorded evidence before seeing the panel’s verdicts.
Delivery belongs in the score when the role requires the specific behavior being judged. Put that behavior in the rubric and ask for evidence under conditions that resemble the work. General ease or composure shouldn’t become a hidden competency.
| What an interview score may reward | What I’d compare instead |
|---|---|
| A strong overall impression | Evidence tied to each question |
| Smooth, assured delivery | The content of the response |
| A broad trait label | The example behind the label |
| Different standards by interviewer | One rubric for each question |
| A group verdict formed in discussion | Individual ratings set before the debrief |
This is the comparison I’d stand behind: shared questions that surface the same kind of evidence, answers rated against each question’s rubric, and a human responsible for the call. The score means the same thing for every candidate the team chooses to assess.
Metaview’s data covers what happened inside the hiring process. It holds nothing about how anyone did once they started the job.
How Metaview makes that comparison.
Metaview Screening uses an AI agent to hold a two-directional screening conversation with each candidate a team chooses to assess. The agent asks questions from the assessment plan and probes when an answer is thin. It evaluates responses against per-question rubrics.
Each rubric shows what the team is evaluating and how it is weighted. Questions can adapt to the candidate’s response while the criteria stay visible.
The output includes an assessed-fit rating with an explanation grounded in the screening conversation, plus full notes and a transcript. A human recruiter decides whether the candidate progresses or is rejected.
Teams can link the call to an applicant tracking system (ATS) stage or share its link individually. Metaview can sync the call summary and the candidate’s information. It can also sync the progress or reject decision once a human recruiter makes it.
The assessed-fit rating leaves a reasoning trail another person can inspect. Metaview’s chief executive officer, Siadhal Magos, describes why he wants that reasoning preserved.
The great thing about using AI for these tasks is you can literally audit the reasoning trail and understand why certain applicants are getting through and why others are not. Once people start to internalize the value of an audit trail on AI reasoning, the bias conversation will totally flip on its head.”
For live interviews, Metaview generates a scorecard from the conversation. The interviewer reviews it, sets the rating, and submits it. Metaview Assistant can answer questions about captured interviews and return the exact details.
Cockroach Labs reports one operational result in its customer story. Lynette Estrada, Vice President of Global Recruiting, says: “With Metaview, our recruiting team has saved over 14 full work weeks.”
Elsewhere in the hiring process, fillmore researches prospective candidates and builds target lists from a role description or hiring brief, writes personalized outreach, books screening calls with interested candidates, and works from Slack.
Confident candidates do tend to outscore less confident ones. A Head of Talent should read that as a warning about what the interview rewards. It doesn’t show that confidence beats competence, so every score should point back to a response and rubric.
Bring Metaview into your hiring stack.
Live notes, structured scorecards, and ATS sync.
Frequently asked.
Do candidates need an account for a Metaview Screening call?
Candidates can take the call on desktop or mobile without creating an account or logging in, whenever suits them.
Can Metaview flag answers that may be fraudulent?
During a screening call, the agent can flag a response likely to be AI-assisted or fraudulent for a human to assess. The flag doesn’t reject anyone.
Can hiring managers compare notes from several interviews?
Metaview can combine interviews from the same hiring process with the resume and job description in one multi-source notes document.
Do Screening questions change based on a candidate’s answer?
Questions can adapt to the candidate’s response while each question’s rubric shows what the team is evaluating and how it is weighted.
Who submits an AI-generated interview scorecard?
The interviewer reviews the generated scorecard, sets the rating, and submits it.