I spend a lot of my week working alongside AI agents, and the thing that surprises people is how much the results vary by who's driving. Hand the same agent to two engineers and one gets a finished feature back while the other gets a confident mess. The gap between someone who can direct an agent well and someone who can't has started to look less like a skill gap and more like a seniority gap.
Companies have noticed. Shopify told its teams that reflexive AI use is now a baseline expectation, and plenty of others have quietly added some version of it to their hiring bar. The trouble is that “works well with AI” has become a line on the scorecard that almost nobody knows how to actually interview for, so it gets waved through on a confident answer and a tool name.
Working with an AI agent isn't really a tool skill. It's a management skill: scoping the work, handing it off, checking what comes back, and knowing when the agent is confidently wrong. You already know how to interview for that when the report is a person. This is how to do it when the report is an agent, with a four-level rubric and the questions that surface each level.
Why working with agents is a hiring bar
Two people with the same title and the same resume now produce very different output, and a lot of the difference comes down to how well they work with AI. One delegates a chunk of the work to an agent, checks it, and ships more. The other does it all by hand, or worse, ships whatever the agent handed back without reading it closely. Over a quarter, that compounds into a real gap in what each of them gets done.
That's why teams have stopped treating this as a nice-to-have and started writing it into the bar. The instinct is right, and it leads somewhere uncomfortable: if you're going to screen for it, you have to be able to tell the levels apart, and most interviews can't. It helps to be clear-eyed about what an agent does, and doesn't do, to a person's output.
If the muscle is atrophied, it will enhance an atrophied muscle. If the muscle is strong, it will enhance the strong muscle.”
That's the whole case for interviewing carefully. An agent multiplies whatever judgment the person already has. Give a strong operator an agent and they get more done with sharper calls about what to trust. Give it to someone without that judgment and you've automated their blind spots. So you're not really hiring for AI use. You're hiring for the judgment the agent will amplify.
The tools your team already touches are turning into agents too. Our own product is built as a set of agents that work across hiring, and the same rule holds there: they reward the people who know how to direct them.
Test for the workflow, not the tool
The usual interview for this is one question, “do you use AI?”, and a nod at the answer. That sorts almost nobody. A candidate can name three agents and still not be able to get good work out of one, and the most confident answer in the room is often the emptiest. Here is the difference between the version that sorts people and the version that doesn't.
- Ask “do you use AI?” and take a yes at face value.
- Reward tool names and the right buzzwords.
- Score from a confident answer in the moment.
- Treat it as one checkbox, same bar for every role.
- Ask for one real piece of work they handed to an agent.
- Probe how they scoped it and what they checked.
- Score the workflow and the judgment, not the vocabulary.
- Grade it on a level, calibrated to the role.
The trap is rewarding the person who is most enthusiastic about AI. Enthusiasm is easy to perform, and judgment isn't. The teams that get this wrong end up hiring the candidate who was busiest with AI rather than the one who got the most out of it, which is its own version of an old measurement problem:
Hiring managers conflate activity with progress.”
Lots of AI activity isn't the same as real progress with it. The best operators draw the line between using AI to go faster and using it to do better work, which is the distinction worth listening for in an interview:
The four-level delegation rubric
A rubric beats a vibe. Four levels do most of the work, and they describe how a person works with an agent rather than which apps they've opened. Read across the row to place a candidate: how they work with agents, what you'll hear when they're at that level, and what kind of role it clears.
| Level | How they work with agents | What you’ll hear | The bar this clears |
|---|---|---|---|
| 1. Operator | Does the work themselves. Uses AI as autocomplete at most, and rarely hands off a whole task. | “I tried it, but it’s faster to just do it myself.” Names a tool they dropped after a week. | Fine where the job is hands-on and AI is a minor convenience. A flag to probe for any role where throughput matters. |
| 2. Delegator | Hands discrete, well-defined tasks to an agent and takes the output mostly as-is. | “I had it draft the first version, then I cleaned it up.” Describes one-off tasks and light checking. | A reasonable baseline for most roles today. Strong on its own where speed helps and the stakes on any single output are low. |
| 3. Supervisor | Scopes multi-step work, sets the context and guardrails, reviews what comes back, and knows where it breaks. | “I gave it the brief and examples, caught where it drifted, and changed how I prompt for that.” Talks about verifying. | Strong for any role where an agent changes throughput. The judgment to catch bad output is the line above Delegator. |
| 4. Orchestrator | Runs several agents like a small team. Builds repeatable workflows others use, and raises the whole team’s output. | “I built the workflow the rest of the team now runs on.” Thinks in systems and second-order effects. | The bar for roles you’re betting the function on. Rare, and worth holding out for when the job is to change how a team works. |
The jump that matters sits between Delegator and Supervisor, and it isn't about tools. It's oversight: whether the person treats the agent as confident and often wrong, and builds that assumption into how they work. A delegator trusts the output. A supervisor checks it on purpose, and can tell you exactly where it tends to go wrong.
The rubric only works if your questions surface real behavior instead of rehearsed opinions. Three moves do most of that:
- Ask for one handoff. “Walk me through the last real piece of work you handed to an agent, end to end. What did you hand off, and what did you keep?” A lived handoff exposes the level; a hypothetical just rewards whoever talks about AI best.
- Inspect the guardrails. How did they scope it, what context did they give, and what did they check before trusting the result? A delegator stops at “it gave me a draft.” A supervisor can narrate how they kept control.
- Probe the catch. “When did an agent do something confidently wrong, and how did you notice?” The answer separates someone who trusts the output from someone who supervises it.
Write down what a strong, average, and weak answer looks like for each level before the interview, and assign the questions to one interviewer rather than all of them. Spread across a panel, this turns into four people asking “so, do you use AI?” and nobody going deep.
Score it from the room, not from memory
A rubric is only as good as the evidence behind each score. The usual failure is that the interviewer asks a sharp question, hears a great answer, and then writes the scorecard two days later from a fuzzy memory of it. The level rounds up, and the rubric becomes decoration.
This is where capture earns its place. Because Notetaker captures every spoken word, the candidate's real example, the work they scoped and the failure they caught, is on the record instead of in your head. The scorecard then drafts from the conversation against the rubric you set, so the delegation level is filled in from what the candidate described, ready for you to confirm or adjust.
From there, Reports shows whether interviewers assessed the competency where they were supposed to, and how the level breaks down across the pipeline and by team. A rubric that lives only in one interviewer's head drifts. One you can see across every loop holds.
It also keeps interviewers honest with each other. You can see where a strong yes from one person means something different from another's, and calibrate before that gap quietly costs you a hire.
None of this replaces your judgment. It gives the judgment something solid to stand on. And the reason it's worth the effort is that the teams treating AI as core to how they work are the ones pulling ahead on the numbers that matter:
Those figures come from Metaview’s 2026 AI & Hiring Alignment Report, surveying 505 recruiting leaders and hiring managers across North America and EMEA. Hiring people who can get real work out of an agent is the other half of that same advantage.
What this means for your team
Pick one role you're hiring for right now and decide, out loud with the hiring manager, which level it genuinely needs. A junior or highly specialized role might only need a delegator. A role you're betting the function on needs a supervisor or an orchestrator. Writing that line down is most of the work, and it's the part teams keep skipping.
Then put the rubric where the interview happens. Add the questions to that loop, build the bar into your question bank and scorecard templates, and let the capture layer record the answers so the score is evidence, not recall. Connect it through native integrations and keep the ATS you already run. If you want the groundwork first, our writeups on great interviewers, quality of hire, and interview quality set it up, Reports keeps the bar honest across the team, and pricing shows what it costs.
Working with AI agents stopped being optional faster than most hiring processes adapted. You don't need a perfect science to keep up. You need a shared definition of the levels, a few questions that surface them, and a way to score that doesn't depend on who was in the room. That's how you interview for the skill that's quietly separating your strongest people from everyone else.
Interview for it, and score it from the room.
Capture the answer, place the candidate on a level against your own rubric, and keep the bar consistent across every interviewer and team.
Frequently asked questions
What does it mean to interview for working with AI agents?
It means assessing how well a candidate can hand real work to an AI agent and get a good result: how they scope the task, what guardrails they set, how they check the output, and whether they catch the agent when it's confidently wrong. It's closer to a management skill than a tool skill, so you interview for it the way you would for delegation and judgment, not for whether someone has used a chatbot.
What are the levels of working with AI agents?
Four levels cover most roles. An operator does the work themselves and rarely hands it off. A delegator hands discrete tasks to an agent and takes the output mostly as-is. A supervisor scopes multi-step work, sets guardrails, checks the result, and knows where the agent breaks. An orchestrator runs several agents like a team and builds workflows others rely on. The biggest jump is from delegator to supervisor, and it comes down to oversight.
How do you interview for AI delegation skills?
Ask the candidate to walk you through one real piece of work they handed to an AI agent, end to end. Then inspect the guardrails, meaning how they scoped it, what context they gave, and what they checked, and probe the catch, meaning a time the agent was confidently wrong and how they noticed. Lived workflows reveal the level; opinions about AI just reward whoever talks about it best.
Should every role require working with AI agents?
Every role now deserves a deliberate line, but the level should match the job. A junior or highly specialized role might only need a delegator, while a role you're betting the function on may need a supervisor or an orchestrator. The mistake is leaving it undefined, so it either gets waved through on confidence or held against candidates inconsistently.
How does Metaview help assess working with AI agents?
Metaview records the full interview, so the candidate's real example and the failure they caught are on the record rather than in your memory. It drafts the scorecard from the conversation against the rubric you set, so the delegation level is scored from what was said in the room. And its reporting shows whether interviewers assessed the competency across the pipeline and how the call lands by team, which keeps the bar consistent.