hiring

I Loved That Candidate. He Wasn't Real.

·8 min read

He was my favorite candidate in the loop.

Great energy. Quick. I’d ask a gnarly system-design question and he’d come back at the pace you’d expect from an opinionated, top-tier engineer who’d done real time at a couple of FAANG companies. No fumbling, no “let me think about that” — just confident, well-structured answers, one after another. The kind of interview where you’re already mentally drafting the offer.

He had some internet issues. Froze a couple of times, dropped audio, came back. Remote life, I thought. Happens to everyone.

It was the freezing that finally nagged at me. Not because it was rude — because the cadence was strange. The connection would stutter, and then the answer would arrive perfectly formed, like it had been waiting. I called the recruiter after and asked what felt, at the time, like a dumb question:

“Hey — did we actually verify he worked at those companies?”

We had not. And when we looked, the story came apart. The energy I loved wasn’t his. The pace I admired was a tell.

I’ve hired a lot of people — operations folks, product managers, engineers, and lately people building AI itself. I’ve been doing this long enough to think I had a decent radar. That one got past me clean, and it rearranged how I think about the whole exercise.

This is not a weird one-off anymore

I went looking, half-hoping I’d been uniquely careless. I had not.

Gartner now predicts that by 2028, one in four candidate profiles worldwide will be fake (HR Dive). In their 2025 survey of 3,000 candidates, 6% admitted to interview fraud — posing as someone else, or having someone else pose as them (Gartner). Six percent will admit it on a survey. Sit with that number for a second.

And the rest aren’t lying so much as leveling up. About 4 in 10 candidates use AI somewhere in the application, and roughly 22% now use AI live, during the interview itself (Newsweek). There are tools — Final Round AI, Interview Copilot — with over a million users that listen through the mic and feed answers onto a second screen in real time. That’s my “great energy, perfect cadence after a freeze,” explained. Harvard Business Review ran a whole piece last fall titled, fittingly, “Are You Interviewing a Candidate — or Their AI?”

Then there’s the truly cinematic end of it. The security firm KnowBe4 hired a North Korean operative who used a stolen U.S. identity and an AI-enhanced photo, cleared four video interviews, and started loading malware on day one (KnowBe4). Four video interviews. These are careful, well-run companies.

So the easy reaction is: lock it down. And 72% of recruiting leaders are doing exactly that — dragging interviews back in person (Newsweek). Which brings me to the part nobody likes to talk about.

The fairness vise

Here’s what makes this genuinely hard, and not just a security problem.

I’ve hired people who don’t speak English as their first language and are brilliant. I’ve hired people who don’t hear or see well, who interview with captions, screen readers, or a little extra time. My whole job is to set those people up to succeed — to make sure the interview measures whether they can do the work, not how fast they can perform under a webcam.

And almost every “anti-cheating” reflex collides head-on with that goal.

Crack down on response latency? You just penalized the candidate on a bad rural connection, the one reading captions, the one translating in their head before they answer. The data here is damning: speech-recognition and AI assessment tools perform measurably worse for people with accents, deaf speakers, and non-native English speakers, and the EEOC and DOJ have formally warned employers that these tools can violate the ADA (ADA.gov). There’s already an ACLU complaint alleging an AI hiring system discriminated against deaf and non-white applicants (HR Dive). When auditors tested these systems, they flagged bias against women, people of color, neurodivergent candidates, and non-native speakers — the exact people I’m trying to give a fair shot.

So that’s the vise. The standard interview can no longer tell a fake from a star. And the obvious “fix” — detect the AI, punish the tell — quietly discriminates against the candidates who most deserve a level field. You can catch the fraud or you can be fair, and the usual tools make you choose.

I didn’t love either option. So I tried something that, when I described it to a friend, made him laugh.

“Ha ha — if it don’t sound crazy, it wouldn’t be worth trying.”

I simulated $1,000 of interviews so I wouldn’t have to burn the credits

The experiment: instead of guessing, I built a little simulation. I had top AI models — Gemini, Claude, others — role-play candidates across the roles I actually hire for (operations, product, marketing, sales, security, engineering, leadership), then had a panel of models score the answers against the same structured rubric a human panel would use. Thousands of simulated interviews. AI answering, AI grading, me reading the results.

Two findings reorganized my thinking.

One: the standard behavioral interview is saturated. The AI candidates scored ~4.8 out of 5. Not because they were caught being robotic — because the answers were genuinely excellent. STAR structure: perfect. “Tell me about a time you showed leadership”: flawless, every time, for every role. If everyone scores a 9, the test scores nothing. The format I’d trusted for years no longer separates people. My fake favorite wasn’t an anomaly; he was the logical endpoint.

Two — and this is the hopeful part — the follow-up is where it breaks down, but only the right kind of follow-up. The generic probes (“what was the result?”, “what would you do differently?”) didn’t dent the AI at all; the answers stayed polished straight through. What did pull them apart were specific, subtle, content-based probes:

  • Counterfactual inversion: “What would have had to be true for you to make the opposite call?” Real experience holds the tradeoff in both hands. A generated answer is confidently one-directional.
  • Numerical reconciliation: accept their metric, then later ask for the number underneath it. “You said 38% — off what base? what’s the absolute count?” Lived numbers reconcile. Borrowed ones drift.
  • Lived, incidental detail: “Take me back to the moment it went sideways — what was the first signal you noticed?” Genuine memory has texture. Fabrication is generic.
  • Role muscle-memory: the exact command an engineer ran, the literal sentence a salesperson used to handle the objection, the precise metric definition a PM measured.

Notice what those probes have in common: they test the substance of someone’s experience, not the speed or polish of their delivery. They don’t care about your accent, your connection, your hearing, or whether you needed a beat to translate the question. They reward you for having actually done the thing — which is the only thing I ever wanted to measure. The faker can’t keep the details consistent; the brilliant non-native-English engineer absolutely can.

That’s the whole point. The fair probe and the fraud-resistant probe turn out to be the same probe. You don’t have to choose between the two sides of the vise — you just have to stop scoring performance and start scoring depth.

Where this goes

I’m turning it into something practical: a way for a hiring manager to keep their standardized questions — which is what keeps the process fair and comparable in the first place — and get a set of subtle, role- and level-tailored follow-ups to ask. The probe you’d use on an IC (“what was the exact flag?”) isn’t the one you’d use on a Director (“walk me through the exact words of the hardest sentence you had to say”). Same question, different depth.

One firm line, for me: this is not an AI lie-detector, and I won’t let it become one. The minute you start flagging “this person sounds like AI,” you’re back to punishing the accent and the lag and the assistive tech. The tool surfaces signal — depth of real experience — and a human reads the substance. Never the cadence. Never the accent.

I still think about that candidate. He was the best interview I ran that quarter, and he didn’t exist. The lesson wasn’t “trust people less.” It was that I’d been measuring the wrong thing — and that the version of the interview that’s hardest to fake is also, conveniently, the version that’s fairest to the real people on the other side of the camera.

If that sounds a little crazy: good. That’s usually the sign it’s worth trying.

This is part one of two. Part two — AI Shouldn’t Pick Your Candidate. It Should Check Their Story. — is on what I went looking for next, and the legal line that decides whether you can use AI to help at all.


Sources: Gartner via HR Dive · Gartner newsroom · Newsweek: AI in interviews · HBR: Are you interviewing a candidate or their AI? · KnowBe4: North Korean fake IT worker · ADA.gov: AI & disability discrimination in hiring · HR Dive: ACLU ADA complaint