Most candidates say they're good with AI, and most of them believe it. The problem is that self-rated AI ability has almost no relationship with performance, and AI tools are designed to feel good to use, so confidence builds regardless of whether someone is catching errors, asking the right questions, or understanding the risks.
AI readiness breaks into six measurable competencies: understanding how the tools work and their limits; spotting the appropriate use cases; identifying bias and risk; knowing what good and bad data does to an output; reading outputs critically; and knowing when a human needs to stay in the loop. Candidates can be strong on some and weak on others, and once you name the criteria you can score them consistently.
Imagine five people on a shortlist, all strong. Somewhere in the process, whether the CV, the screening call, or the "tell me about a time you used AI" question every guide now includes, all five have told you they're good with AI. They’re comfortable with it, they use it every day.
What makes this hard is that most of them aren't exaggerating to impress you, they believe it. And a self-assessment someone believes is a difficult thing to argue with, unless you have a way to measure the claim fairly, repeatably and objectively.
Mistaking confidence for signal
"I'm good with AI" answers a different question from the one you're asking. You want to know whether someone understands these tools well enough to use them properly: whether they can tell when an output is wrong, judge when it's safe to rely on, and know where the risks are. What they're telling you is how they feel about doing that.
Think about how someone comes to believe they're good with AI. They open a chatbot, ask it something, and get a fluent, confident answer... it feels great. But the people you want are the ones who catch that the answer was subtly wrong, or that they asked a question the tool was always going to flatter. The tool is built to feel good to use, and that feeling arrives whether or not you're any good at it.
There's evidence for this. A peer-reviewed study out of Monash last year measured people's GenAI ability directly and compared it against how good they said they were. The self-ratings were no real predictor of performance; the direct measure was. How good someone says they are at AI tells you very little about how good they are.
The cost of not really measuring AI readiness
The obvious cost of a weak signal is bad hires, and the development spend that follows when you can't trace poor performance back to the screening question that let the wrong person through.
But the biggest cost is the reason you were hiring for this at all. You wanted AI-ready people to get something out of AI: faster work, better output, a team that can use the tools you're paying for. Depending on whose research you read, 80% to 95% of enterprise AI efforts deliver no measurable return, and RAND puts AI project failure at roughly double the rate of ordinary IT projects. Most of that is data, strategy and governance rather than people. But whether your workforce can work with these tools is one of the few factors within your control at the point of hire. The real cost isn't a slightly worse hire. It's that the person who could have moved the needle got hired by your competitor instead, because you couldn't tell them apart from the one who talked a good talk.
And the interview won't save you. Interviews reward the articulate and composed, and being articulate about AI is a completely different skill from being good at it. You can talk well about a tool you barely understand, the tool will even help you prepare!
What you need to be measuring
"Good with AI" resists definition because it stands in for several things. Break it down and it's six distinct facets you can define and assess:
- How the tools work and their limits (AI Foundations);
- Spotting a real use case rather than forcing AI everywhere (Application and Use Cases);
- Seeing where it goes wrong, from bias to harm (Bias, Ethics and Risk);
- What good and bad data does to an output, and where sensitive data shouldn't go (Data Quality and Governance);
- Reading an output critically instead of letting a fluent answer end the thinking (Output Interpretation and Evaluation);
- Knowing when a human has to stay in the loop (Human Oversight and Collaboration).
A candidate can be strong on some and blank on others, and that's the point. Once you've named the criteria, you can tell whether a candidate has each one, and rank two candidates on the same basis instead of on impression.
Where that leaves you
The candidates aren't the problem. The confident ones might even actually be your strongest hires. The problem is you can't tell, because you're screening on how they feel about AI, not what they can do with it.
The good news is that it's fixable. AI readiness breaks down into things you can define and score, the same way you already assess reasoning or judgement. And the same measure does more than improve hiring: you can run it on your existing teams and it shows you where to point development, so training targets real, measurable gaps instead of assumptions. The useful next step is to look at where "good with AI" is already influencing who you hire, and ask whether you'd be able to defend your decision under scrutiny. If you would, carry on. If you wouldn't, you already know the signal needs replacing, and now you know it can be.




.webp)
.webp)
.webp)

.webp)
