A Recruiter's Guide to Vetting in the AI Era

html<table style="border-collapse: collapse; font-size: 13px; width: 100%; margin: 0 auto;">
  <thead>
    <tr>
      <th style="border: 1px solid black; padding: 4px 6px;"></th>
      <th style="border: 1px solid black; padding: 4px 6px;">Delta TPs</th>
      <th style="border: 1px solid black; padding: 4px 6px;">Full TPs</th>
      <th style="border: 1px solid black; padding: 4px 6px;">Total TPs</th>
      <th style="border: 1px solid black; padding: 4px 6px;">Likely FPs</th>
      <th style="border: 1px solid black; padding: 4px 6px;">Likely FP Rate</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="border: 1px solid black; padding: 4px 6px;">Claude<br>Code</td>
      <td style="border: 1px solid black; padding: 4px 6px;">44 / 46<br>(95.7%)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">19 / 50<br>(38.0%)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">62 / 95<br>(65.3%)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">48</td>
      <td style="border: 1px solid black; padding: 4px 6px;">43.6%</td>
    </tr>
    <tr>
      <td style="border: 1px solid black; padding: 4px 6px;">Codex<br>(GPT-5.5)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">43 / 45<br>(95.6%)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">30 / 50<br>(60.0%)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">74 / 95<br>(77.9%)</td>
      <td style="border: 1px solid black; padding: 4px 6px;">629</td>
      <td style="border: 1px solid black; padding: 4px 6px;">89.5%</td>
    </tr>
  </tbody>
</table>
<p style="font-size: 12px; font-style: italic; margin-top: 8px;">Table 2: True positive (TP) and false positive (FP) analysis of Claude and Codex across challenge types.</p>
Table of contents

Every hire is a trust decision made on incomplete information. You read some claims. You ask some questions. You make a decision to either hire someone or pass on them. That’s always been the case.

What has changed in the last 18 months or so is that it now takes seconds to manufacture a convincing, yet inaccurate, resume thanks to AI, while the time it takes to verify it hasn’t dropped at all. Many think the answer is to simply filter harder. It isn’t. To understand why filtering fails, we need to look closer at the problem. Below, we’ll cover:

  • Why hiring operates as a reputation system
  • Three common types of manufactured identities we’re seeing today
  • What every phase of an interview loop should actually verify
  • Better strategic alternatives to aggressive filtering

For an engineering role we recently opened, we got 1,800 applications in the first 24 hours. Let’s be clear. We are in an incredibly competitive job market. We aren’t dealing with a stream of applicants. It’s a flood, and somewhere inside that deluge of applications are the people who are exactly who they say they are. I gave a presentation about this at BSides Las Vegas in August, and it’s worth highlighting here.

Hiring is a Reputation System

Resumes, certifications, endorsements, and LinkedIn profiles only reflect a candidate's claimed identity. It is the version of their background they choose to present, without any independent verification. Nobody at career or networking sites has checked any of it, yet it’s still treated as verified evidence. 

The risk is clear: Companies grant real access based on unverified claims. It starts with an interview, escalates to a job offer, and ends with access to your codebase and customer data.

Every reputation system has the same failure mode. If a career or networking site can’t tell the difference between true and fabricated experiences, whoever can manufacture their resume fastest wins. There are many accounts claiming employment at places they’ve never set foot in, and asking LinkedIn to manually verify or take them down just isn't realistic.

There’s a name for that. In 1973 there was a book called Sybil, written by Flora Rheta Schreiber. It tells the story of a woman  who presented as 16 different identities. Microsoft researcher Brian Zill borrowed the name from the famous book, and it quickly became a standard term in computer science.

A Sybil attack is when an adversary floods a system with fake identities, and if the trust system can’t distinguish between them, one attacker gains the influence of many. 

Any hiring pipeline is a trust system, and it is under exactly this kind of attack from both directions. Candidates flood it with inflated identities. Managers answer with more filters. The filters punish the honest people and train everyone else to fake harder.

Anatomy of a Fake Identity

The word “fake” makes this sound simpler than it is. Manufactured identities are a layered problem. The three versions I run into most often are separated by exactly one thing: how deep into your process they can get before anything gives them away.

Here are the three most common patterns I see: 

Inflated Identity: The resume said “led.” The reality? They only “assisted.” The use of misleading vocabulary is rampant, because it’s now the cheapest thing in the world to fake. And short introductory calls with a recruiter, known as phone screens, mostly test vocabulary, which is why this identity clears them so reliably. I’ve watched a candidate sail through a take-home assessment and a hiring manager interview round before collapsing 45 minutes into a two-hour panel as soon as we asked them to go deep.

Proxy Identity: As a recruiter or hiring manager, you authenticated the face of the person being interviewed, but do you actually know if the person you hired is the one doing the work? Sure, the candidate passed the interview, but they could outsource the actual job to a third party for a cut of their pay. That leaves your systems exposed to people you have never vetted. You tend to find out this has happened when an unfamiliar name shows up on a file or in a codebase. 

Multiplexed Identity: One resume. One interview. One badge. The problem is, though, the hands on the keyboard change day-to-day. This is essentially the DPRK IT-worker playbook: treating a single identity as shared infrastructure. Your first stage in the recruiting process should catch this. The problem is that deepfakes are getting good enough that it increasingly doesn’t. 

Notice what these three issues have in common? None of them should survive verification. All of them survive a process that only checks where the claim is well-informed.

Every Stage has to Verify Something

Think of the interview loop as a protocol. It has two lanes: candidate and company. A handshake happens in stages. Every stage exists to verify one specific thing.

Stage

Protocol analog

Verifies

Resume review

SYN

Well-formed claims

Recruiter screening call

Challenge-response

Consistency, plus one layer of depth

Hiring manager screening call

Key exchange

Real material, both directions

Technical

Proof of work

Thinking, not performing

Final panel

Certificate signing

The team stakes its reputation on yours

The only job of the resume review stage is to be well-formed enough to earn a response. The recruiter screen tests consistency and looks for enough depth to confirm there’s a real operator behind the resume before it costs an engineer an hour of their valuable time. The hiring manager’s screen is where both sides reveal real material. The technical screen is to get proof of work, and that is where inflated identities die. The final panel is certificate signing. Your team is putting its own reputation on this person.

Your hiring loop probably doesn’t look exactly like this. Plenty of companies fold the hiring manager screen and the technical round into a single conversation. They also might split the technical screen into two parts. That’s fine. The stage names are the part you can move around. Meaningful verification is the part that has to be done consistently.

Now try writing that third column for your own process. One sentence per stage: this stage verifies X.

If you can’t write the sentence, you don’t have a protocol. You have vibes with extra steps. Any stage that verifies nothing is noise in your own pipeline. And that costs you and the candidates time, while adding zero information to the decision process.

Why Filtering Harder is Failing

The instinct right now is to filter harder. That instinct is wrong, and it’s the reason we’re here.

An applicant tracking system (ATS) will hand you an AI match score for every applicant, and I have watched hiring managers open only the resumes that scored 100 and never look at the ones that scored a 75. Sometimes, the 75 is the better candidate. The system misses that because it measures keyword overlap. It does not measure judgement, and judgement is the entire thing you were trying to hire for.

Filtering harder just means that candidates will work harder to optimize against the filter. That’s a rational response to a system that discards the un-optimized version. The system taught people to run Sybil attacks.

Verification is the opposite move. It costs more per candidate but it works because the thing it tests can’t be generated. Ask the candidate about a problem they had to solve, never the tool they used.  And follow up until you hit the edge of what someone knows. Real experience has friction in it, meaning candidates will talk about dead ends, things that broke, or maybe decisions they still aren’t sure about. A clean, uninterrupted victory from start to finish is likely a highlight reel of somebody else’s work.

The same logic also runs the other way. If you’re the candidate, the most valuable thing you own is work a third-party can authenticate. That means pointing to concrete proof, like a CVE, a writeup, a repo with real commit history, a talk, or a CTF placement. Public work is signed work. It doesn't have to be elite. It has to be checkable. A blog post about a bug you genuinely struggled with beats a claim of mastery every time, because one of them can be checked and the other can’t.

Proof Beats Credibility Every Time

Hiring is a reputation system, and reputation systems only survive if they can verify.

Every reputation system in history that lost the ability to verify eventually got Sybil’d. Our hiring pipelines are not special. The flood isn’t going to stop. The tools are only going to keep getting better, and no filter you build is going to outrun a generator.

So stop asking whether a claim looks credible and start asking what would verify it. That’s a harder question, and it’s the only one that still has an answer.

The real ones are unmistakably real once you know what you’re checking for.

By clicking Sign Up you're confirming that you agree with our Terms and Conditions.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.