When a loop screens for fundamentals and then walks each candidate through a take-home they built at home, it measures whether the candidate can produce working code.

The failure it leaves room for comes after the offer. It is a change that passes review and continuous integration and still breaks something in production, and neither the reviewer nor the engineer who shipped it can explain why the code was written that way.

Nothing in that loop tests the vibe coding competency. It never measures whether the candidate can stand behind code they did not personally write, which on a team that works this way can be most of what they ship.

That gap has a name now. Vibe coding is a working practice, and the useful response is to specify what it requires.

I would stop asking whether a candidate wrote the code, and start scoring what they did to make it safe to ship.

The technical interview is the longest stage in the loop, averaging 62.6 minutes, which makes it the one conversation with room to ask how the code got there.

This post sets out what the vibe coding competency is, and how to set a bar for it that your interviewers can apply in the loop you already run.

Screening for authorship is the wrong test.

Loops that adapt to artificial intelligence (AI) defensively add a proctored round, or they ban assistants outright.

Some set a puzzle novel enough that a model hasn’t seen it. All three answer the same question: did this person write the code themselves?

On the job, the answer is increasingly no, and that is sanctioned. An interview that screens for unassisted authorship selects for a working mode the role doesn’t use.

The question worth an hour of a hiring manager’s time is narrower and harder. When an engineer ships code they didn’t write line by line, what did they do to make it safe, and what did they do to make it possible for the next person to own?

Those behaviors are observable, and they vary enormously between candidates. They also predict what happens after the offer.

The four parts of the vibe coding competency.

Break the competency into parts you can score independently, because candidates are rarely strong or weak across all of them. The common profile is fluent generation with thin verification, and that profile passes most interviews.

Specification quality.

The spec is now the engineering artifact. A vague prompt produces plausible code that solves a slightly different problem, and the cost of that mismatch surfaces late.

Strong candidates describe constraints before the answer: what the input really looks like in production, and what the code must never do.

Verification against intent.

The failure mode here is specific and common. The candidate writes the tests after the fact, reading them off the generated code, so what the tests lock in is whatever the model happened to assume.

Those tests pass every time, and they say nothing about whether the requirement was met.

Reading code they didn’t write.

Review, debugging, and explanation are now the bulk of the job, and they’re harder on unfamiliar code than on your own.

What you’re looking for is whether the candidate can walk a path through the code they shipped and name its failure modes without rereading it from scratch. Generating quickly and reading carefully are different skills, and a candidate can be strong at one and weak at the other.

Judgment about blast radius.

This is the part almost no loop assesses, and the part that separates a senior engineer from a fast one.

Vibe coding is appropriate in some places and reckless in others, and the boundary has little to do with technical difficulty. It is set by how far the damage travels and how long it takes to undo.

A candidate who applies the same method to a throwaway script and to a migration hasn’t yet developed the judgment the role needs.

Where the bar sits, by blast radius.

A single standard for the whole codebase is what makes these conversations go badly. Ban the practice and you lose speed everywhere it was safe.

Wave it through and you get the change that passed review and still broke production. The bar moves with blast radius, and saying so explicitly is a hiring manager’s call.

Code area What goes wrong unsupervised What you require before it ships
Prototype or spike Very little. The risk is that it becomes permanent without anyone deciding that it should. A stated expiry and no route to production. Speed is the whole point here, so don’t tax it.
Internal tooling It breaks in six months and the person who generated it cannot debug it. A named owner who can explain it cold, plus enough of a test to fail loudly.
Customer-facing product code Tests written from the generated code, which confirm the bug they were meant to catch. Tests derived from the requirement before the implementation exists, and a reviewer who has read the whole diff.
Auth, billing, migrations Failures here are silent and expensive. Once data has moved, they are hard to reverse at all. Line-by-line review by someone who could have derived the change themselves, before anything merges.

Publishing that table internally does two things at once. It gives your interviewers a shared standard to score against, and it gives candidates a fair account of how the team works.

That account is worth more in an offer conversation than any assurance about culture.

Questions that surface real judgment.

The questions below assume the candidate used AI and move straight to what they did about it.

Phrased as an accusation, the same question gets a defensive answer and nothing you can score. Ask about a specific artifact they shipped, because a philosophy of AI is easy to talk about and hard to check.

  • Walk me through something you shipped where a model wrote most of it. What did you check before it merged, and what did you decide not to check?
  • Show me a test you wrote for that change. What would have to break for it to fail?
  • Where in your current codebase would you refuse to work this way? Listen for a specific subsystem and reason. A general caution doesn’t count.
  • Tell me about a time the output looked right and wasn’t. How long before you noticed, and what changed in your process afterwards?
  • Someone else has to own this code in six months. Ask what they left behind for the next owner, and how they know it is enough.

The last two do most of the work. A candidate who has been burned describes the delay before they noticed, and that detail is difficult to invent convincingly.

For the failure-and-boundary line of questioning, and the trap of reading too much into a single good answer, there’s more in interviewing for AI judgment.

Running the assessment in interviews you already have.

None of this needs a new interview stage. It needs the existing technical conversation to capture more than a verdict, because the evidence for this competency is in what the candidate said.

Most written feedback does not carry that much. Across 13,382 written scorecard feedback entries, the median entry runs 365 words, and 26.6% of them stop at five words or fewer.

26.6%
of written interview feedback runs to five words or fewer, against a median of 365 words across the same sample.Source: Metaview research, 13,382 written scorecard feedback entries, data pulled June 2026

The median shows the standard is reachable inside the interview teams already run. A five-word entry is not always a weak judgment, but it cannot carry the walkthrough the judgment came from.

Metaview’s Notetaker captures every spoken word of the technical conversation, including the parts an interviewer would not think to write down, so the scorecard draft starts from what the candidate said.

I don’t think the answer to an information problem is making decisions with even less information.”
Siadhal Magos Siadhal Magos Co-founder and Chief Executive Officer · Metaview

Much of a vibe-coding answer is never spoken. The clip below shows the notetaker picking up code written during the interview, the changes made while debugging, and the system design diagram on screen, alongside the transcript.

So the trace-a-path answer and the what-would-this-test-catch answer are still there when you fill in the scorecard, an hour after the call.

Metaview Notetaker showing AI notes alongside the live interview transcript
The walkthrough is captured as it happens, so the reasoning is there to review later.

Scorecards then carry the four parts as four named criteria. Scoring specification, verification, reading, and blast radius separately is what stops a fluent generator from clearing the bar on fluency alone.

It also means two interviewers scoring the same candidate are scoring the same four things.

Metaview scorecard fields populated from the interview transcript
1
2
3
  1. 1Each competency is pre-filled from the transcript.
  2. 2Every rating carries its evidence and the timestamps it came from, so a reviewer can check the claim.
  3. 3The suggestion is a starting point for the interviewer, who edits and submits it.
Scorecard criteria populate from what was said in the interview, with the transcript behind each one.

At SoSafe, Director of Talent Acquisition Fiona Keating says the team uses Metaview for “stronger hiring signals, better calibration across interviewers, and a really high-quality experience for candidates as well, regardless of the interviewer.”

Score the vibe-coding criteria from the interview itself.
See how Metaview drafts the scorecard from a technical interview, for the interviewer to review and submit.
See it live

Reports is where a hiring manager finds out whether the new criteria are being used at all.

A competency that appears on the scorecard and never in the conversation is a competency no one is assessing, and that gap is visible across interviewers well before it shows up in a bad hire.

Metaview Reports showing competency coverage across the interview pipeline
Competency coverage across the pipeline shows which criteria are being probed.

Pick one engineering role and add the four criteria to its scorecard for the next loop.

Publish the blast-radius table alongside them, so every interviewer is calibrated against the same standard.

After a handful of interviews, look at which criteria are being skipped. The ones that never get probed are the ones to fix first.

Usually that is verification and blast radius, because they are the two that require the interviewer to slow down.

See it in action.

Put the four vibe-coding criteria on your next engineering loop.

Score specification, verification, reading, and blast radius from what the candidate said.

Frequently asked.

Should candidates be allowed to use AI during a coding interview?

Allow it wherever the role allows it, which for most engineering roles it now does. Say so in the invite, so candidates are not left guessing, and tell them up front that they will be asked to explain and defend whatever the tools produce.

How do you brief an interviewer who has never assessed this?

Give them the blast-radius table and one worked example of a strong and a weak answer for each criterion. Most interviewers can judge specification and code reading on instinct, because both resemble what they already do in review. Verification and blast radius are the two that need a written standard before the first interview.

What do you do when two interviewers score blast-radius judgment differently?

Treat it as a calibration problem before you treat it as a candidate problem. Go back to what the candidate said about the subsystem they would refuse to touch, and check whether both interviewers were working from the same definition of expensive to undo. Disagreement that survives that check is usually a real split in the team’s own standard, and it is worth settling before the next loop.

How do you assess a candidate who prefers to work without AI?

Score the same four criteria against the code they did write. Specification, verification, reading unfamiliar code, and judgment about what a change can break all predate these tools. A candidate who is strong on all four is well placed to apply them to generated code once your conventions are clear.

Does this apply to junior hires, or only to senior engineers?

It applies to both, at different weights. For a junior hire, specification and verification decide most of the score, and blast radius is something the team teaches in the first few months. For a senior hire, blast radius is the criterion to weight most heavily, because its absence goes unnoticed until an incident.