How to Test Whether Candidates Can Spot Mistakes in AI-Generated Answers (an Interview Round You Can Run)

By Pooja Goenka ยท 2026-10-07

Imagine a candidate reading five lines an AI wrote about Python default arguments. The answer is fluent and mostly right. She nods and says, "Looks good to me." Line five is wrong. You never asked her to write code, and you've just learned something a coding test would have missed.

Short answer: give the candidate a confident, well-structured AI answer with exactly one planted technical error, ask them to review it out loud, and score whether they find the line, explain why it's wrong, and leave the correct lines alone. A shared document and a scorecard are enough. Below are the steps, a worked example, the scoring, and the tools that can help.

Why this skill matters now

A Sonar survey of more than 1,100 developers, published in January 2026, found that developers estimated AI wrote 42 percent of the code they committed. Of them, 96 percent didn't fully trust it to work correctly, only 48 percent said they always verify it before committing, and 38 percent said reviewing AI code takes more effort than reviewing a colleague's.

IEEE Spectrum reported in September 2026 on teams coping with the volume. Synthesia's CTO said pull requests were up 120 percent year over year as of August, and 95 percent contained AI-generated code. Bonterra's CTO said proposed changes tripled within three months of adopting AI and review times tripled too. Temporal's "Send Back" policy rejects agent-written code unless the engineer explains, in their own words, the agent's design choices and how the code handles unusual conditions.

A June 2026 arXiv paper by Vahid Garousi argues that oversight of AI output is mandatory: engineers must review, validate and sometimes rework what the AI produces. It's a synthesis of eight sources, seven of them practitioner writing, so it's a framing, not a measurement.

You'll also see claims that engineers now spend more hours reviewing than writing. The only figures I found came from a post that didn't name who ran the survey, so I've left them out. The sources above support something narrower: AI output needs a human check, and many developers admit they don't always do it.

For interviewers, "can you write this function?" is easy to answer with an assistant open. "Can you tell when this confident answer is wrong?" is harder to fake, and it's the job.

The round, step by step

It takes about ten minutes, as a screening call or a block inside a longer interview.

What you hand the candidate

Give them a question they could meet in the role, then an AI-style answer of five to seven numbered lines. Numbers let them point. Follow these rules:

  • Exactly one line is wrong. If you can't defend that, rewrite it.
  • The wrong line is a claim a busy person would accept, such as a convenient rule or a detail that's almost right.
  • Every other line is correct, and at least one correct line sounds fussy or surprising. That's your decoy.
  • All lines share the same tone and length, with no hedging.
  • You've checked every line yourself. An AI can draft the answer, but it sometimes plants a mistake you didn't ask for.

How long, and what to say

Allow one minute for the brief, three for a silent read, four for discussion and two for probes. Say this:

"An AI wrote this answer. It sounds confident, and one line contains a technical mistake. Review it the way you'd review a colleague's pull request. Think out loud, and tell me which line you'd send back and why."

Then stop talking. Don't say "good" or "close" while they work.

How to push back without giving the answer

React the same way whether the candidate is right or wrong. If they approve everything, or pick a line, use these probes in order:

  1. "How sure are you? What would you check before approving this?"
  2. "Pick the line you trust least. How would you test it?"
  3. "Say this ships and breaks in production. Where do you look first?"

If they attack a correct line, ask, "What would you run to check that?" Strong candidates often talk themselves out of it. Never name the line.

What to write on the scorecard

Record the line they flagged first and the line they ended on, the reason in their own words, how many probes you used, whether they proposed a check (a test, a doc lookup, a quick run), and whether they criticised a correct line and how confidently.

A worked example

The question: a teammate's function keeps remembering items between calls, and the signature is def add(item, bucket=[]). Why?

  1. Python evaluates default argument values once, when the def statement runs, and not on each call.
  2. So bucket=[] creates a single list that every call without that argument shares.
  3. Items added in one call therefore show up in later calls.
  4. The usual fix is to default to None and create the list inside the function with if bucket is None: bucket = [].
  5. This only affects lists. A dict or a set as a default is safe, because Python makes a fresh copy of those on each call.

Line 5 is the mistake. Python copies nothing. Dicts and sets are mutable too, so a {} or set() default is created once and shared exactly like the list. I ran all three in Python 3 and each carried data from the first call into the second.

A good reason: "Dicts and sets are mutable and created once at definition time, so they share state across calls just like the list." It names the mechanism. A weak reason: "dicts and sets are slower than lists, so they shouldn't be defaults." The candidate may have found the line by instinct, but the reason has nothing to do with the bug.

Lines 1 and 4 are the decoys, and both are true. Line 1 trips people who half-remember the rule and "correct" it to "evaluated on every call", which is backwards. A candidate who confidently attacks either has flagged a correct line, which is a different result from missing line 5. Write it down separately.

How to score it

Use a 0 to 3 scale and keep the notes beside it.

What happened What it tells you Suggested score
Caught the line and gave the right reason Reads each claim and understands the mechanism 3
Caught the line, reason missing or wrong Good instinct, may be pattern matching 2
Flagged a correct line, missed the real one Confident criticism with no checking 0
Missed it, approved everything Trusts fluent output 0

If a candidate needed a probe, drop them one level and note which probe. If they say "I can't tell for sure, here's the test I'd write", give that credit in the notes even when the line was wrong. A confident wrong critique still scores as a miss, however articulate.

Common mistakes interviewers make

  • Choosing trivia. If the error depends on an obscure flag, you're testing memory. Pick something the candidate would meet on the job.
  • Planting two errors by accident. Test every line, and have a colleague read it cold.
  • Leaking through reactions. A nod or "mm-hm" near the right line gives it away.
  • Rewarding the longest critique. Score the line and the reason, whatever the length.
  • Reusing one answer for every candidate. It will spread, so keep a few variants.

Tools that can run or support this

You don't need a tool for the version above. Here's what the main platforms say on their own pages, read on 7 October 2026.

HackerRank describes Chakra as an AI interviewer. A HackerRank post from 5 August 2026 says AI fluency is tested with scenarios such as reviewing AI-suggested code, catching a subtle bug a model introduced, and deciding whether a generated solution is good enough or needs a rewrite. Scoring runs across several dimensions, not a single pass or fail. TechCrunch reported on 5 October that Chakra gives candidates real repositories with an AI assistant, asks follow-up questions about their reasoning, and has run more than 500,000 interviews.

Codility's page on AI in technical assessment asks whether an engineer can judge if AI-generated code is correct, efficient and appropriate. It looks at whether a candidate engages with AI output critically or accepts the first answer, and it describes AI-generated follow-up questions that check whether candidates understand their own solution.

CoderPad's AI Fluency docs say interviewers can watch how a candidate prompts, iterates, evaluates AI output and applies judgment, and that reviewers should check whether the candidate verified the output. The AI Fluency page is aimed at non-technical candidates. The docs give little detail on scoring criteria.

CodeSignal's blog post describes an assistant called Cosmo and gives hiring teams a transcript of the candidate's AI interactions plus a session replay, which shows whether candidates ask informed questions, iterate well or lean on suggestions. It doesn't describe a planted-error review.

Placed.dev lists a GenAI coding round where the AI gives you code that looks right but isn't, and you find the bug, fix it and explain the change. MockIF describes a debugging round with a planted bug and a voice AI interviewer who scores your process. The MockIF page doesn't mention reviewing AI-written answers.

CloudApper's AI Recruiter page describes scenario-based questions and configurable scoring, and its article on AI in interviews is about detecting live AI help. None of the CloudApper pages I opened describes giving a candidate AI output to review. Kalpita's homepage lists Kalpita AI Interview as an enterprise AI product with no detail on how it assesses candidates, so I can't say how it handles this.

LogicWiz Interviews is one more option. Its inverted round has the AI write a confident, well-structured answer with a subtle technical error planted in it and ask the candidate to review it. If they miss the error, the AI pushes back gently without giving it away. A candidate who confidently criticises the wrong thing scores as a miss, and a structured scorecard records the result. Candidates can practise free with a free account: 3 AI sessions a day, no card, 16 tracks and 100 questions, by voice or text. There's also a free five-minute game, Spot Aria's Mistake, where you tap the wrong line across five rounds and then pick why. The question bank is small, and LogicWiz isn't a full coding-environment assessment tool. If candidates need to work in a repository, the platforms above suit that better.

Common questions

What is an inverted interview round?

The AI or interviewer supplies the answer and the candidate does the checking. The answer sounds confident and contains one planted error.

How many errors should the answer contain?

One. A single error keeps the scoring clean and lets you tell a miss from a wrong critique. Add more lines, not more mistakes.

Should I tell candidates there's a mistake?

For a ten-minute round, yes. Without the hint you mostly measure suspicion. For senior roles you can drop it and score whether they ask what to verify.

Can I use ChatGPT or another AI to write the answer?

For a draft, yes, but check every line yourself. An AI asked for a subtle mistake sometimes plants two, or one that's actually correct.

Is this round enough to make a hiring decision?

No. It shows how someone reviews and verifies. Pair it with a build or debug task.

Sources I read

All opened and read on 7 October 2026.