Amazon's bar raiser is the most copied hiring idea in tech. Teams copy the ritual: a senior stranger with a veto, dropped into the loop near the end. They skip the reason the role exists at all, which is that somebody in the room has to be free of the pressure to fill the seat.
There's an awkward problem with writing about bar raisers, so let's get it out of the way first. The whole idea is named after a promise about who you end up hiring, and that promise can't be checked.
Interview records end at the offer. Whatever happened after the person started sits in a different system, owned by a different function, on a different cadence, and it rarely gets joined back to the interview it came from.
So when you read that a bar raiser program lifts the standard of the people you hire, you're reading a claim with no data behind it. That includes us.
The process is where the answerable questions live. Did a standard exist in writing before the loop opened, did every interviewer apply it, and did the decision leave a record?
That's a smaller claim than the name promises, and it's the one worth making. The reason to run a bar raiser is that the role puts one person in the loop whose defined job is to leave behind an account of the decision a stranger could read.
Judge the program on that account, because that's the part you can actually inspect.
What a bar raiser is, minus the mystique
Amazon's version is an interviewer from a different team, trained for the role, who joins the loop and carries weight in the final call. Everything interesting about the design sits in who they answer to.
A hiring manager owns a requisition and a date. A bar raiser owns neither of those, so one person in the room has no reason to go along when the search drags on and the standard starts sliding.
That's the part worth copying, and it costs nothing. You don't need a training pipeline or tens of thousands of interviews a year to drop an interviewer from another team into a loop and give them standing in the debrief. Two or three of them will cover a small org.
Those numbers are a starting shape and nothing more. No dataset anywhere tells you the correct ratio of raisers to loops, and anyone who quotes you one is guessing.
Most teams copy the veto instead, which is the least important part of the whole design. Give someone the power to block and nothing else, and people learn to route around them. Make the loop answer an objection in writing and you change how the loop runs, whether or not the raiser blocks anything.
The claim a bar raiser program cannot make
The name carries the claim: raise the bar, and the people you hire get better. Put that claim aside, because no recruiting team can check it, and the vendors selling you software can't check it either.
Interview data stops at the decision. Metaview Reports covers what happened inside the recruiting process, and it holds no dimension at all for what came after the offer. A better query won't close that gap.
Whether someone worked out gets recorded on a different cycle, by different people, against criteria the interview never used.
The closest thing to an outcome measure anyone has here is self-reported. In Metaview's 2026 AI and Hiring Alignment Report, 79% of teams with an excellent recruiter-to-hiring-manager relationship and high alignment exceeded their business goals, against 36% of teams whose relationship was fair or poor and whose alignment was low.
Read what that actually measures. Two survey answers, cross-tabulated against a third, from 505 recruiting leaders and hiring managers. It's an association between reported relationship quality and reported goal attainment.
The survey didn't test interviewer consistency and it didn't test bar raiser programs, and with no causal design it can't tell you which way the arrow runs.
So the case for a bar raiser gets made on the process or it doesn't get made at all. That turns out to be fine, because the process is where the visible problem sits.
What the interview record does show
Start with what a loop leaves behind. Most loops leave less than you'd guess.
A scorecard was attached to 31.2% of the interviews in a corpus of 5.2 million candidate interviews. That's the normal state across the interviews in this corpus.
It shifts once the first version writes itself. In a capped sample of 120,000 scorecards, 50.3% of the AI-generated ones got submitted (n=93,502) against 28.6% of the manually written ones (n=26,498).
The 31.2% and the 50.3% count different things and don't belong in one series. The first is coverage across interviews. The second is submission across scorecards created.
The same gap shows up at the moment it matters most. Take 296,555 advancing candidates: 41.9% of them had no submitted scorecard at the point they moved forward. They advanced anyway, and the reasoning behind it never got written down.
What most loops lack is a record of the judgment anyone can read afterward. Put one person in the loop whose explicit job is to leave that record behind and you've fixed a common failure cheaply.
The tiebreaker story doesn't hold up
The tempting version of this argument says rooms are quietly split and the raiser is there to break the tie. The record doesn't support it.
More than one person scored the same candidate in 72,753 loops, and 85.0% of those ended unanimous. Most rooms already agree, so a second reader is rarely deciding anything.
The loops that do split are the more interesting ones. At least one dissenting scorecard sat on file against 56.3% of the candidates who advanced from a loop where more than one person scored them.
That denominator conditions on the candidate advancing, which over-samples exactly the loops where somebody objected, so it isn't evidence that dissent gets routinely overruled. What it does tell you is that a recorded objection and a halted process are two separate events.
So here's what the raiser is actually for: making sure the objection that already sits in the record gets an answer in the record. Breaking ties was never the job.
There were six items hiring managers were supposed to assess for, but they all had their own questions. We'd never calibrated on what a good answer looked like.”
That's the precondition, and it's rarer than it sounds. The standard usually lives in a structured interview guide. A raiser with no written standard behind them is one more opinion in the room, arriving late.
How to stand up a lightweight program
Here's a program shape that holds up at small scale, and the written standard and the independence carry all the weight. Headcount is the easy bit.
- Pick two or three raisers. Go for the interviewers whose read you'd trust on a role outside their own. They keep their day jobs, and bar raising is a hat they put on for other teams' loops.
- Calibrate on real interviews. Have them score the same set of past interviews and compare, before they raise the bar for anyone. Your standard is vague wherever they diverge, and sorting that out in a room together is the training. Work from a shared question library so everyone starts from the same material.
- Write the standard down before anyone interviews. Say in plain language what clears the bar for each role family. A standard that lives in one person's head can't be taught, can't be applied by anyone else, and can't be defended to the candidate you turned down. Our writeup on interview scorecards covers the format.
- Staff one raiser per loop, from outside the team. Give them standing in the debrief. They don't need a unilateral veto, and handing them one usually backfires: define the role by blocking and people learn to schedule around it.
- Decide in advance where an objection goes. Agree that a no on the loop gets a written answer before the candidate moves, naming what the objection was and what changed the room's mind. Teams skip this step, and it's the only one that shows up later in the record.
We wanted very clear and objective criteria so that we could measure every candidate equitably and consistently. We needed to make sure that those areas were evaluated thoroughly across the board without any unintentional or unconscious bias creeping in.”
What to check once it is running
Here are four questions you can answer from the process itself, each one with a limit attached. The limits matter as much as the questions. A program that only reports the left-hand column is telling you it worked without having tested anything.
| What to check | Where the answer lives | What it still does not tell you |
|---|---|---|
| Did a written standard exist before the first interview? | The dated guide or scorecard template for the role family. | Whether the standard is the right one. A written bar can be written badly. |
| Did every interviewer on the loop submit a scorecard? | Submission per loop, counted against the interviewers who were scheduled. | Whether the scorecards say anything. A submitted skeleton still counts as submitted. |
| Did the debrief turn on what the candidate said? | The interview record, sitting next to the notes it came from. | Whether the room read the evidence correctly. Two people can quote the same answer and disagree. |
| When someone scored no, was it answered in writing? | The loop's scorecards and the written decision that followed them. | Whether the objection was right. An answered objection is documented rather than settled. |
Answering those checks is mostly a capture problem. Metaview's Notetaker joins the interview as a visible participant with the candidate's consent and records the conversation, so you review what was actually said instead of what somebody remembered two days later.
Metaview Reports queries your own interview data across the team, which is what lets you answer the first two rows across a whole quarter of loops. The judgment stays exactly where it was: with the raiser, in the debrief, against a standard a human wrote.
One more check is worth running at the role level instead of the loop level, and that's coverage. Did the same competencies get probed for the same role in March as in January?
You don't need to replace anything to run this. Keep your ATS and connect Metaview through native integrations. The customers page has the detail if you want to see how other teams run their loops, and our writeup on what separates great interviewers is a reasonable place to start on the training side.
What this means for talent leaders
Run the program with what you already have: two calibrated raisers, one written standard per role family, one raiser in every loop from outside the team, and an agreement about where an objection goes. It takes a quarter to stand up and it costs you no headcount.
Then be careful about how you report it. The temptation is to reach for the outcome story once the program is running, because that's the story executives want to hear.
You can't support it, and reaching for it is how a decent process control gets oversold and then quietly dropped the moment someone asks for evidence. If someone asks whether the program changed how the hires turned out, tell them that measurement doesn't exist, here or anywhere, and that a bar raiser program was never going to settle it.
Report what's real and specific instead: every loop opened against a written bar, every interviewer's read on file, every objection answered on the record. That record is what a bar raiser produces, and it's what to judge the program on.
Give your raisers something to read.
Capture every interview, keep the scorecards and the objections in one place, and let the debrief turn on what the candidate actually said.
Further reading: interviewer training is the same discipline, one level down.
Frequently asked questions
What is a bar raiser in hiring?
A bar raiser is an interviewer who sits on a loop for a team other than their own and holds the standard for the role rather than the deadline. The name and the training pipeline behind it are Amazon's, where the stated aim was that every hire should raise the average. Everywhere else it's a hat rather than a job: the person keeps their day role and raises the bar on a handful of loops a quarter. Pick someone who can judge the work rather than someone in recruiting, because the raiser has to be able to tell a strong answer from a fluent one.
Does a bar raiser program improve who you hire?
There's no data that can settle it: interview records end at the offer, and post-hire outcomes live in a different system that's rarely joined back to the interview. If an executive asks what the program returned, the answer worth giving is a process one, and it's stronger than it sounds. You can show that a written standard existed before the loop opened, and you can show what every interviewer wrote against it. Promising more than that is how the program gets cut the first time somebody audits the claim.
Do you need to be Amazon to run a bar raiser program?
No, and the size question matters less than the ownership question. Someone has to own the standard, keep the raiser list current, and notice when a loop ran without one. That owner is usually recruiting rather than any single hiring manager, because a hiring manager owning it recreates the conflict the role exists to remove. Rotate the raisers on a schedule you set in advance: a raiser who sits on the same team's loops all year stops being a stranger to that team's habits, which was most of the point of having them there.
What does a bar raiser actually do in an interview?
They run their share of the loop, then do the part that's theirs alone in the debrief: press on which competencies were genuinely covered and on what the candidate said, rather than on how the room feels. The question worth settling before it comes up is what happens when the raiser and the hiring manager deadlock. The workable version is that the hire can proceed over an objection, as long as the person proceeding writes down what the objection was and what changed their mind. That decision then sits in the record with a name against it, which is a heavier thing to write than most people expect.
How do you start a lightweight bar raiser program?
Start with calibration rather than with the org chart. Put your candidate raisers in a room with the same set of past interviews, have them score independently, then compare. The disagreements are the syllabus, and one session rarely settles them, so plan for a few sessions over a few weeks rather than a single afternoon that ends in an untested agreement. Only then write the standard down for each role family, and only then put a raiser in a live loop. Standing the whole thing up takes about a quarter, and most of that is calendar time on the calibration sessions.
How do you keep the hiring bar consistent across interviewers?
Consistency decays, so treat it as maintenance rather than as setup. Re-run the calibration exercise when the role changes, and again when a standard has gone a year without anyone arguing about it, which usually means it stopped being read. Keep the interview record so a loop can be reviewed after the fact rather than reconstructed from memory. And leave a route open for the standard itself to be wrong, because none of the process checks tell you the bar is set in the right place, and the raisers are the people most likely to notice when it isn't.
