Asking a candidate to answer a screening question in their second language forces them to do two things at once: think through the answer, and translate it in their head.
That takes longer, and it shows. The answer can come out slower, shorter, or plainer than what they know. If your rubric rewards how an answer sounds, that can look like a weaker candidate, whether a recruiter or an artificial intelligence (AI) screening agent does the scoring.
Multilingual AI screening can take some of that pressure off, when your team offers the screen in the candidate’s strongest language and holds every version to one rubric. I’m for it, as long as it’s done right. It’s still new, so test it on your roles first, and buy from a vendor you’d trust with your candidates.
This guide covers when a candidate who doesn’t speak your language is still the right hire, how to write a rubric that holds up across languages, and when to get a second read from someone who speaks the candidate’s language.
Should a candidate who doesn’t speak your language still get the job?
Often, yes. If they’ll work in your team’s language day to day, which is usually English, they’ll need a working knowledge of it. But working knowledge isn’t fluency.
Someone who reads your docs a little slower, or takes a beat longer to reply in a meeting, can still be the strongest person for the role.
Three things tend to arrive as one requirement: the language the work happens in, the language the team uses, and the language the hiring manager would prefer. Pull them apart.
Start with the work. A support role that answers customers in Spanish needs spoken Spanish. It might only need enough English to read internal tickets, a much lower bar than taking calls in English.
So agree on both with the hiring manager when the role opens, and say how good each needs to be. “Can handle a customer call in Spanish” is something a screen can test. “Speaks Spanish” isn’t.
The team’s language counts when the new hire will use it every day. Even then, the bar is usually working knowledge, and if the work itself happens in another language, the team’s language may only be a preference.
Then there’s the hiring manager’s preference. Wanting a new hire you can chat with in your own language is understandable. But in that support role, the best candidate might be the one answering in their second language.
So talk the preference through when the role opens. Ask what the job loses without it, and keep it out of the rubric.
If no one says it out loud, it can creep in through rubric words like “clear” and “polished,” which reward how smoothly someone talks. That’s a subtle kind of screening bias.
A language need is one more criterion, so decide which stage tests it, as you would on a hiring map.
Write the rubric around what the answer shows.
A rubric band that describes how someone sounds scores their language along with the job. It does that whether a person or a tool scores it, and whoever wrote the rubric.
A band that describes what the candidate did or decided holds up much better across languages. Here are four rewrites:
| The band says | What it rewards in a second language | A band that names the evidence |
|---|---|---|
| “Communicates clearly and confidently” | How fluent the candidate sounds | Explains the steps they took, in order, and why they chose them |
| “Articulate, polished answers” | Vocabulary and register | Gives one specific example and names what came of it |
| “Uses the right terminology” | Knowing the term in the screen’s language | Describes the process correctly, in their own words |
| “Great rapport with customers” | Accent and pace | Resolves the customer’s problem in the role-play and states the next step |
Then check each language version. It’s your team’s question set in that language, and each question has to ask for the same evidence in every version. Translation can shift a question in small ways:
- An idiom with no match in the other language.
- A level of formality that changes what counts as polite.
- A word that narrows the question.
- An example that only makes sense in one market.
So before the screen goes live, have someone fluent read each version against the rubric. Note who checked it, and when.
Follow-up questions asked live never get that check. That’s one reason some results need a second read later.
Brex ran into a version of this in its sales hiring, where mock demos were hard to judge. In its case study’s words, “things like energy or style can easily overshadow substance.”
Its team wrote structured rubrics into Metaview interview templates and used Metaview reporting to coach interviewers. “That early 20 percentage point jump in onsite-to-offer rates tells us we’re on the right track,” said Danielle Harders, Brex’s director of global business recruiting.
That’s one team’s early result, from interviews, and the rubrics came in alongside reporting and coaching. So it can’t show what a rubric does alone, and it says nothing about language.
I’d still expect the habit to carry over: write down what good looks like before anyone scores. The Brex story also runs as a two-minute film on Metaview’s channel.
Decide how fluency and language switching count before anyone is invited.
My rule of thumb: when a role doesn’t need a particular language, let every candidate answer in their strongest language, from the ones your team can read. When it does, test it openly.
Write both rules down before the first invitation goes out. They’re your team’s rules, whichever tool runs the screen.
The first is for a language the job needs. Give it its own question. Ask it in that language and score it at the level the job needs, like handling a customer call. Tell candidates which question is in which language, and why.
A stumble on that question tells you something about the language. It shouldn’t drag down their other answers.
The second covers everything else. On every other question, no band scores fluency, register, or accent. Switching languages mid-answer, or borrowing an English technical term, should cost the candidate nothing. Write that into the rubric.
Then list the languages candidates can answer in, keeping to the ones your team can read. The detected language mix across conversations captured on Metaview shows how long that list can get.
Get a second read when language might have moved a rating.
Send a result to a fluent recruiter, one who reads the candidate’s language, when:
- The answer is in a language the reviewing recruiter doesn’t read.
- The reason behind a rating mentions fluency, register, accent, or vocabulary, and the role doesn’t need the language.
- The candidate switched languages partway through an answer.
- An answer is so short that language, rather than knowledge, may be holding them back.
Interviewers disagree often, and that’s normal. Of the candidates who advanced with more than one submitted scorecard in interviews captured on Metaview, 56.3% had at least one scorecard recommending against them.¹
I’d settle any disagreement the same way: go back to what the candidate said. If they said it in Portuguese, that means someone who reads Portuguese.
Metaview Screening supports 18 languages and scores each response against that question’s rubric, which your team can edit. It gives the recruiter the reason behind each rating.
The recording keeps each answer in the language the candidate spoke. With the transcript laid out question by question, a fluent recruiter can check a rating against what the candidate said.
For every completed call, a recruiter decides who moves forward, and Metaview does not make that decision.
The sample reason in the screenshot above ends on a small deduction for “a slightly formal register in speech.” For a retail role where a warm tone with customers is part of the job, that may be fair.
Where it isn’t, a band that rewards register is scoring how someone speaks the language. Settle it in the rubric, then use the written reason to check the rating followed it.
Good interviewers know they are contributing to a hiring decision. Bad interviewers believe they are making a hiring decision.”
So make the second read part of the process. Name a reader for each language, and have them send a short note with the result to the recruiter who decides.
Four checks that show whether multilingual AI screening holds one standard.
Put these on next quarter’s review:
- Advance rate by answer language. Is one language well below the role’s overall rate, with enough candidates to tell? Check its questions first, before the candidates or the pool.
- What the lowest ratings lost points for. If fluency, register, accent, or vocabulary shows up for a role that doesn’t need the language, rewrite the band that let it in.
- How often the second read changed the call. If it never does, narrow the triggers. If it often does, something is scoring language, so check the bands first.
- When each language version was last checked. If it was before the rubric last changed, check it again before the next invitation goes out.
Start with the second one. The reasons behind the lowest ratings are the likeliest place for a language preference to show up in writing. Fixing the band behind it is the first step back to one standard.
See the evidence behind every screening rating.
Walk through a completed screening call, from the rating to the answers behind it.
Frequently asked.
What if a candidate answers in the team’s language with a strong accent?
Score the answer, and leave the accent out of every band the job doesn’t need. For English in the United States, the EEOC’s guidance allows a decision based on an accent only with evidence that effective spoken English is required for the job and that the accent “interferes materially with job performance.” If a rating’s reason mentions accent, someone listens to the recording before the recruiter acts on it.
Who reads the answers if no one on the team speaks the language?
Someone elsewhere in the company who does, briefed on the rubric ahead of time. If no one in the company reads it, leave that language off the invitation for now, or the recruiter ends up acting on a rating no one can check.
Should screening questions be translated or written fresh for each language?
Start from the evidence each question should draw out, and have a fluent speaker write the question to ask for that. A literal translation keeps the words but can change the question.
What if the job needs two languages?
Then both are criteria, each with its own question and its own band. The rest of the answers can come in whichever language the candidate is strongest in.
Does any of this change if a person runs the screen instead?
Mostly no. An AI screen adds one more reader between the answer and the rating, which is why the second read goes back to the recording. If a recruiter runs the call on Metaview’s Notetaker, which captures every spoken word, you can check what each candidate lost points for against the call itself. Without that record, the recruiter has to write down why each candidate lost points.
Sources.
¹ Aggregated and anonymized Metaview interview data, all-time: of the 130,818 candidates who advanced with more than one submitted scorecard, 73,688 had at least one “No” or “Strong No” recommendation. No individual, company, or candidate is identifiable.