34.9% of engineering job applications get flagged medium or high risk by Metaview's fraud-detection model. That's the highest of any function, ahead of data at 32.8% and well clear of sales, product and design.
It's worth being precise about what that measures. The model scores applications at the point of submission, looking for identity deception and automated applying. All of it happens at the top of the funnel, before anyone opens an exercise.
What it tells you is that engineering funnels arrive noisier than anyone else's. Every filter you put downstream is carrying more weight than it did two years ago.
Which is why take-home versus live coding stopped being a scheduling preference. Here's what the data says about which one to use, and what to change about both.
Where the AI problem actually sits
Ask a technical hiring manager what AI broke and most will say the take-home. A candidate gets a repo and 48 hours, and a model that writes competent code in 90 seconds sits in another tab.
That's a fair description of the take-home and a bad description of the problem's size. The same model sits in another tab during your live round too, or on a second monitor, or on a phone propped against a laptop stand. Screen share only proves what's on one screen.
Neither format was ever verification. A take-home never proved the candidate worked alone, and a live round never proved they weren't reading from something off-screen.
What changed is that the distance between a strong engineer's output and a weak one's narrowed. The artifact carries less information than it used to. What didn't narrow is the gap in how they explain the work.
Technical interviews are already your most structured stage
Engineering teams get criticized for interviewing badly. The data doesn't really support that, at least not next to everyone else.
68.8% of technical interviews have a scorecard template attached before the call. For a generic interview it's 31.7%, and for the intake call between recruiter and hiring manager it's 0.4%. Technical rounds also run the longest, averaging 62.6 minutes against 35.3 for intake.
So the technical stage is where most companies already get it right. There's an agreed rubric, it's settled before the call, and there's an hour of real work to score against it.
That moves the diagnosis. If your technical loop is producing bad hires, the likelier culprits sit either side of it: the intake call where requirements were never written down, and the debrief with nothing to work from.
Every single candidate at Vercel has to have a work product for us to submit an offer. It can vary, some take-home exercises, some presentations, some specific exercises live with the interviewer. There is something we submit for all candidates.”
Vercel's answer to the format question is that it isn't one. Every candidate produces something, the mechanism varies by role, and what's fixed is that a work product gets attached to the decision.
The question most interviews still skip
Here's the gap that surprised us most. Only 15% of interviews touch the candidate's use of AI tools at all. Six in seven technical conversations never get to it.
The trend is moving fast in relative terms. AI usage as an interview topic went from 0.33% of sessions in Q3 2025 to 4.54% by Q2 2026. That's still nowhere near where you'd expect for work these tools have reshaped.
This is the cheapest change available to most teams, and it needs no new tooling. If you think a candidate leaned on a model for the take-home, ask them how.
Ask what they prompted, what came back wrong, what they threw away, where they stopped trusting it. An engineer who used Claude well can talk about that for ten minutes and it's genuinely informative. An engineer who pasted the problem in and shipped the first answer cannot.
Zapier published a fluency framework that adapts well to interviews, and we've turned it into an AI fluency rubric. Scoring it turns AI use into a competency you assess rather than a rule you police, which is closer to how the job works.
How to run a take-home worth scoring
The take-home still earns its place. It's the closest thing in your process to the job: real tools, real time to think, no audience.
It rewards people who do good work over people who perform well in a room, and it surfaces technical signal a resume screen misses entirely. What has to change is what you score.
- Cap it hard. Three hours, and say you'll only read three hours of work. Long take-homes select for people with free evenings.
- Start them in a codebase. Greenfield problems are what models are best at. Ask for a change to existing code with an awkward constraint, and the exercise becomes reading comprehension before it becomes typing.
- Allow AI, ask for the record. Banning it is unenforceable and selects for candidates willing to lie. A short note on what they used gives you something to score.
- Never score the artifact alone. The submission buys a 30-minute conversation, and the conversation is what gets scored.
That last one carries the weight. Walk through their choices, then ask for a change they didn't anticipate: a new requirement, a failing edge case, a performance constraint.
Someone who understands their own submission handles that in a few minutes. Someone who doesn't stalls immediately, and it's obvious to everyone on the call.
How to run live coding that tests more than nerves
Live coding gives you what a take-home can't: the reasoning as it happens. You see what a candidate does when they're stuck, whether they ask before assuming, and how they take a correction.
A capable engineer who freezes in front of an audience is still a capable engineer. Teams know this and walk into it anyway, and a round that mostly measures composure will keep telling you quiet people are worse at their jobs.
- Let them use real tools. An IDE, documentation, and yes, an assistant. Then watch how they direct it. A bare text editor tests a setup they'd never work in.
- Pick an ambiguous problem. The best minutes are usually the first five, before any code is written, when a good candidate stops to ask what happens to the empty case.
- Get out of the way. Where we can measure it, interviewers take a median 46.8% of the conversation. In a coding round that's too much.
- Debug something together. Handing over broken code and working through it is a better proxy for the job than any from-scratch problem.
We've written a longer walkthrough of how to lead a coding interview, including the competencies worth naming in advance.

Which one to use, and when
Both formats work. The choice is about which risk you'd rather carry for the role in front of you.
| Situation | Take-home | Live coding |
|---|---|---|
| Senior and staff engineers | Strong. Design judgment needs room a 45-minute round compresses out | Best used as the walkthrough round |
| Early-career hires | Weaker. Output converges, so it separates candidates least where you need it most | Strong. Reasoning is what distinguishes them |
| Candidates holding other offers | Costly. Unpaid hours are the first thing they decline | Strong. One scheduled hour and it's done |
| Pairing and mentoring roles | Blind to it. You never see them work with anyone | Strong. Collaboration is the thing being tested |
| Candidates who interview poorly under observation | Strong. This is the population take-homes exist for | Risky as a sole gate. Pair it with something else |
One number gets misused here. Candidates in interviews of 61 to 90 minutes advance at 30.4%, against 21.5% for 31 to 60 minutes and 14.2% for anything half an hour or under.
It's tempting to read that as longer interviews producing better signal. It almost certainly runs the other way, because longer interviews tend to be later-stage interviews and those candidates have already been filtered. Extending your coding round to 90 minutes won't double anyone's advance rate.
It's simple. We do a cognitive assessment, a cultural assessment, and a coding assessment for engineers. It's a data point in our hiring process. It's not pass-fail. It's a data point. We're looking at the holistic candidacy.”
Where both formats actually break
You can pick the right format, write a good problem, run the round well, and still end up with a decision the panel can't defend a week later. The break happens after the interview ends.
Only 31% of interviews get a scorecard at all. Where one is written, 26.6% of the feedback runs to five words or fewer. And early-round scorecards agree with the final verdict 54.4% of the time.
Those three numbers describe one situation. The evidence from a 62-minute technical round lives in the interviewer's head for a few hours and then becomes an impression.
By the debrief there's a verdict and very little of the reasoning behind it, which is how panels end up agreeing 85% of the time about candidates they can barely describe.

What moves the number is making the record cheaper to produce than to skip. When interviewers start from a blank scorecard, 28.6% get submitted. When the first draft is generated for them, that rises to 50.3%, and those come back with about three times as many fields filled in.
That doesn't prove drafting caused the whole gap. Teams that switch it on may already run tighter processes. But the mechanism is obvious enough: reviewing a draft an hour after a coding round is a different task from rebuilding one from memory.
Vercel's Viet Nguyen went further and used the record as his calibration set. "Doesn't matter if it was a values or coding interview, strong yes meant offer. So I dumped all his transcripts into ChatGPT to extract the actual behaviors that distinguished strong-yes from yes from fail." That only works if the transcripts exist.
One last finding, and it's the least expected here. We looked at which words show up more in positive scorecards than negative ones. The list isn't technical.
Candidates called a good communicator appear 2.09 times more often in a yes than a no. Adaptable 1.96 times, thoughtful 1.83, personable 1.80, pragmatic 1.74. Proficient, the closest thing to a competence word near the top, sits at 1.57.
Read that alongside the format debate and it lands somewhere uncomfortable. Both take-homes and live coding are built to measure whether someone can write the code.
What tips a technical hire from a maybe to a yes is how they work, and that shows up in the conversation around the code rather than in the artifact. Whichever format you pick, the conversation is the part to capture.
All figures here come from Metaview's own platform data, covering 5.5 million interviews captured since 2020. They are observational comparisons across customer teams, so they show association rather than cause.
Bring Metaview into your hiring stack.
Coding and system design rounds route to purpose-built note templates, so the code, the debugging and the diagrams land in the summary your panel works from.
Frequently asked
Is a take-home or a live coding interview better for hiring engineers?
Both are work samples, so both beat abstract puzzles. Live coding suits early-career hires and roles built on pairing, while a take-home suits senior work where design judgment needs room. The stronger predictor in either case is the conversation about the work.
Should candidates be paid for take-home assignments?
If the exercise runs beyond a couple of hours or touches real product work, paying is the fairer call and it improves completion among senior candidates weighing competing processes. A short, capped exercise on a synthetic codebase is generally accepted unpaid.
How do you tell whether a candidate actually understands code they submitted?
Ask them to extend it live with a requirement they didn't anticipate, such as a new edge case or a constraint that breaks their current structure. Someone who wrote and understood the solution adapts it within a few minutes. Someone who didn't will stall on the first change that touches their own design.
Can you run a technical interview without an engineering background?
A recruiter can run the process and score the observable behaviors: how the candidate handles ambiguity, takes a correction, and explains a trade-off. An engineer still has to judge technical depth, which is why the round is worth capturing rather than rebuilding from memory afterwards.
Is whiteboard coding still worth running?
Abstract puzzle questions test puzzle familiarity more than job skill, and strong candidates increasingly decline processes built on them. A short realistic task in the tools the person would actually use gives you more to score.
What should you do when the panel disagrees after a coding round?
Disagreement is more common than most loops admit: 56.3% of candidates who advance out of a multi-interviewer loop had at least one dissenting scorecard. Rather than resolving it in the room, go back to what each interviewer actually observed, because a split usually means two people watched different parts of the same interview.