A rating that can't be traced to what the candidate actually said doesn't survive the first disagreement in a debrief. Whoever gave it ends up defending an impression, and the colleague who scored the candidate differently has nothing concrete to test it against.

Before anyone asks a behavioral question, the panel needs to agree on what a complete answer contains and which parts of it carry the score. A plan that lists the questions and stops there leaves each interviewer to decide that alone, halfway through someone's answer.

Two questions, two kinds of answer.

Two question types get used interchangeably inside the same interview, and they don't do the same job.

A behavioral question asks for something that happened. "Tell me about a time you shipped without the data you wanted." The candidate returns a past episode. Because it happened, it carries details you can push on, like who else was in the room and what the constraint really was.

A situational question asks for something that hasn't. "What would you do if engineering told you the date was slipping?" The candidate returns an intention. There's no record behind it, so nothing in the answer can be checked.

Both belong in an interview. The mistake is scoring them the same way, and that's what happens when the interview plan never labels them. It's the same discipline that makes competency-based interviewing work, applied one level down to the answer rather than the competency.

Interviewing is not a naturally occurring capability. Most people want to believe that their ability to assess others is high quality because it's relevant every single day in every part of their life. They want to believe they have a good read on people. But when they get into an interview context, it's not natural to do a good job without training, thought, and practice."
JM Jordan Mazer Head of Talent at speedrun · Andreessen Horowitz

The four parts of an answer you can score.

Most interviewers have heard of STAR. Fewer use it as a scoring instrument, which is what it was built for. It names four parts of an answer, and a complete answer has all four.

The situation is the context, and you need just enough of it to know what was at stake. The task is what this person was responsible for, which is not the same as what the team was responsible for.

The action is what they personally did. The result is what happened, and how they knew.

Those four are not equal. The action is the part that tells you what this person did rather than what happened around them, so it's the part to score hardest.

Part What a complete answer contains What it sounds like when it's missing
Situation A specific occasion, with enough context to know what was at stake. "We were always dealing with tight deadlines."
Task What this person owned, stated separately from what the team owned. "The team was responsible for the migration."
Action What they personally did, and the choice they made when it wasn't obvious. "We got everyone aligned and moved forward."
Result What happened, and how they knew it had happened. "It went really well and people were happy."

The failure worth naming is the one that looks like success. "We shipped it and the numbers moved" reports an outcome and skips the person. The action is still missing, and it passes unchallenged because it sounds like an accomplishment.

The repair is a narrow question rather than a broader one. Ask what they did. What happened next is a different question. That is one specific case of the wider habit covered in follow-up interview questions.

Scoring a situational answer, which has no result.

When you ask what a candidate would do if engineering told them the date was slipping, you've written the situation for them. What comes back is a plan, and there's no result to report. Scored against the four parts, even a strong answer gets marked down for gaps the question itself created.

What you can score is the reasoning, as long as you've written down beforehand what good reasoning sounds like for that competency. A strong answer brings up a constraint you didn't mention, then commits to one of two options that could both be defended. Weaker answers describe both options and never choose.

Ask what they'd do if the date slipped after the client had already been told. To answer, they have to weigh what the client was promised against what the team can deliver, and that trade-off is the reasoning you're scoring. Ask how they handle missed deadlines in general and you'll get a policy statement instead.

Candidates sometimes answer a hypothetical with something that happened to them. When they do, score the episode against the four parts like any other. It's better evidence than the hypothetical you asked for, even if it doesn't test the trade-off you had in mind.

A situational answer shows how someone reasons through a problem, but it can't show whether they'd follow through. On the same competency, a real episode should always outweigh it.

Separate fluency from evidence.

Candidates are widely coached on STAR too. So a fluent four-part answer does tell you something real, and the first thing it tells you is that this person prepared.

That cuts both ways. A candidate who narrates in the first person may have been coached to, and one who says "we" throughout may just work somewhere that talks that way.

Neither reading is available from the delivery, so the delivery is the wrong place to look. Ask what they would do differently now, and score the answer to that question.

Score what the answer contained. How smoothly it arrived is a separate fact, and a less useful one. Treat fluency as neither credit nor suspicion, and put the weight on the questions that surface evidence.

Keep the answer where the panel can check it.

A method only helps if the answer survives the interview. Answers go missing between the room and the debrief, when whoever heard them is rebuilding the conversation from memory.

Metaview's Notetaker exists for this part of the problem.

Capture the answer while you listen.

Recording runs on consent, and the consent process belongs to your team. What candidates are told reads: "If both you and the interviewer consent, you will see a 'Metaview Notetaker' appear as a participant during the interview."

It records and transcribes. The practical effect for you is that you can follow the action instead of writing down the situation.

Draft the scorecard from what was said.

The Notetaker drafts a scorecard from the conversation. Browser extensions autofill the objective sections for Ashby and Greenhouse, and the subjective fields are left blank on purpose. The rating and the hiring recommendation are yours, and you submit it.

I still need to add my feedback/takeaways to the candidates responses especially what important things that weren't said! I don't need AI trying to infer things from the conversation."
TS Toni Shilling TA Manager · Prokeep

Direct submission into the ATS works for Ashby and Lever. Greenhouse requires the user to paste. You can also choose or switch the notes template that structures the draft.

Metaview notes template picker showing Topic Highlights applied and five additional structures
1
2
3
  1. 1The currently applied notes template is clearly marked.
  2. 2Each template describes what it covers, and the structured interview templates show their section counts.
  3. 3Switching templates rewrites the call notes in the selected structure, with no advance choice required.
Choose a notes template to structure what the draft covers, from question-and-answer detail to role alignment.
Want this set up on your interviews?
Connect Metaview to your ATS.
Book a demo

Put the evidence in front of the panel.

Several conversations plus a resume can be pulled into one document, which is what a debrief needs. The argument then happens over what the candidate said rather than over who remembers it better.

The Calls list shows who ran each interview and which interviewers have already submitted a recommendation.

Metaview Calls list showing past interviews, named interviewers, interviewer-submitted recommendations, and ATS sync status
1
2
3
  1. 1Filters narrow the list by call type, this week, or your own calls.
  2. 2Each row keeps the candidate or call, type, duration, and named interviewer together.
  3. 3The Recommendation column lists the named interviewer's submitted rating, and the ATS column shows sync status.
Past calls, each listed with the interviewer who ran it. Data shown is illustrative.

What a method cannot tell you.

A method makes answers comparable across interviewers.

Nothing here tells you how someone will do once they're hired. The interview record ends at the hiring decision, so it can't speak to what happens after it.

What a shared definition of a complete answer buys you is narrower, and still worth having. Two interviewers who disagree can find out why.

Give the debrief something to point at.

Pick the shape of question you need before the interview, then write down what a complete answer to it contains. Score the parts you heard. You'll leave the room with something a colleague can check.

If you want a starting set of questions to apply this to, the interview questions library is organized by competency. And Metaview's Notetaker is where the answers go once you've asked them.

Metaview Notetaker

Keep the answers where your panel can score them.

See how a candidate's answer gets from the interview to the debrief.

Frequently asked.

How many behavioral interview questions should one interview include?

Fewer than most interviews attempt. Each one needs room to keep asking until you have the action, so a few questions explored properly beat a longer list left at the first answer. Pick the competencies this interview owns and let the rest of the loop cover the others.

What if a candidate's strongest example is confidential?

You don't need the proprietary details to score the answer. Ask what they were responsible for and what they chose between. How they knew it had worked usually arrives in the same breath. The action and the reasoning are what you're scoring, and neither requires naming the client or the numbers.

How do you score an answer when the candidate only describes what the team did?

Ask again, and ask narrowly: what did you decide, and who disagreed with you. If the second answer still comes back in the plural, score the competency on what you have and record that it went untested. An unproven competency is a gap in the interview, and the rest of the loop can close it.

Can you compare answers from interviewers who asked different questions?

Only if they agreed what a strong answer looks like for that competency. The questions can differ. The definition of a strong answer can't, or the two ratings aren't measuring the same thing and the debrief has nothing to resolve.

Sources.

Jordan Mazer, Head of Talent at speedrun, Andreessen Horowitz, on the Metaview podcast.

Toni Shilling, TA Manager, Prokeep, on Metaview's Wall of Love.