Your AI Skills Might Not Be Keeping Up With Your Confidence

AISA participants expected higher AI fluency scores than they received. Other research raises related questions about self-assessment. Here’s how to check whether your confidence is earned.

Written By
Corey Noles
Corey Noles
Sep 15, 2026
6 minute read

Getting a chatbot to produce an impressive answer feels like a win. Deciding whether that answer deserves to leave the chat window takes another kind of skill.

That distinction gets uncomfortable when you ask people how good they think they are at using AI and then assess them.

In an August analysis, AI skills assessment company AISA reported that 274 participants predicted an average score of 63.4 out of 100, then received an average score of 44.9. Their expectations overshot their assessed performance by 18.5 points. Even the engineering subgroup overestimated its scores by 15.4 points, according to AISA’s findings.

The numbers raise a question worth asking before your next AI-assisted assignment: What evidence do you have that you’re using these tools well?

Research beyond AISA suggests that people can benefit from AI while struggling to judge how well they performed with it. That leaves room for a particularly awkward combination: better work, misplaced confidence, and errors nobody thinks to check.

What AISA’s numbers actually measure

AISA’s assessment uses an AI-led conversation and separate AI evaluation to score participants against a rubric. Its published methodology covers prompting, critical thinking, technical understanding, workflow application, and safety.

In the article’s larger dataset of 1,172 assessments, workflow application was the highest-scoring category at 49.6. Technical understanding and safety were lowest at 41.3 and 41.5. Those category scores come from a larger group than the 274 participants with prediction data; they don’t establish a separate confidence gap for each skill. AISA’s analysis

There are limits to what readers should conclude. These are results from an assessment provider’s participants, rather than a representative survey of all professionals. AISA also acknowledges that its format evaluates reasoning about tools rather than live tool execution, samples a single session, and can advantage articulate communicators. Its methodology lists empirical test-retest reliability data as a planned next step.

The defensible takeaway is specific: participants expected to perform considerably better on this assessment than they did. That’s a useful starting point for examining confidence, even if it can’t tell you exactly how capable your coworkers are.

Advertisement

Is this Dunning–Kruger with a chatbot?

It’s tempting to reach for the Dunning–Kruger effect: the familiar idea that people with weaker skills can also lack the insight needed to recognize their shortcomings.

It's interesting and enough to perk your ears up, but an average gap between predicted and assessed scores doesn’t fully demonstrate that pattern. And additional research examining AI assistance directly adds a twist.

In “AI makes you smarter but none the wiser,” researchers investigated performance and self-assessment on logical reasoning tasks. In the first study, 246 participants used AI to answer 20 LSAT reasoning questions. Their performance exceeded a comparison norm by three points, but they overestimated their scores by four.

A second study involving 452 participants replicated the findings. Higher AI literacy correlated with less accurate self-assessment. The researchers also found that the usual Dunning–Kruger pattern disappeared under AI assistance in these tasks.

That doesn’t establish that AI knowledge causes overconfidence, or that every experienced user misjudges their work. It does challenge a comforting assumption: knowing more about the technology doesn’t automatically mean knowing how well you’re performing with it.

AISA and these researchers measured different things. Together, their findings justify asking whether confidence is receiving enough scrutiny alongside capability.

A polished answer can still need work

Another piece of the picture comes from Anthropic’s AI Fluency Index, which examined 9,830 multi-turn Claude conversations during one week in January 2026.

When conversations produced artifacts such as code, documents, or interactive tools, users gave more direction. Yet visible fact-checking was 3.7 percentage points less common, and questioning the model’s reasoning was 3.1 points less common.

Anthropic offers several possible explanations. Finished-looking work might invite less scrutiny. Different tasks might require different kinds of evaluation. Users could also be checking outputs elsewhere, by running code or asking a colleague to review a document.

The study therefore doesn’t prove that attractive formatting switches off critical thinking. It does highlight a distinction worth remembering: giving detailed instructions and evaluating the result are separate activities.

Advertisement

You can specify the audience, tone, structure, and formatting perfectly—and still receive a beautifully organized mistake.

Five ways to check your own confidence

These are practical exercises, rather than a validated test of AI proficiency. The aim is to make your expectations and results easier to compare.

1. Replace “looks good” with a definition of good

"You are done when..." is smart thing to get used to telling your AI. Before asking AI to complete a task, write down what would make the result usable.

For a research brief, that could mean every major claim has an accessible original source, all figures refer to the correct period, and the conclusion follows from the evidence. For a spreadsheet, it might mean totals reconcile and formulas behave correctly when inputs change.

Then assess the output against those requirements.

If your only standard is whether you like the answer, you’ll have trouble distinguishing genuine improvement from better presentation. The Neuron’s guide to defining goals and testing prompts includes exercises for making those checks explicit.

2. Explain one important conclusion without the chatbot

Pick a recommendation or claim from the result. Explain, in your own words, why it holds up.

What evidence supports it? What assumptions does it depend on? What new information would change the conclusion?

You don’t need to reproduce every step yourself to benefit from AI. But if you’re responsible for a recommendation and can’t explain its basis, you’ve found a specific gap to investigate.

Ask the model for an explanation if that helps. Then check the explanation, too.

3. Give verification somewhere independent to happen

Asking a chatbot to review its answer can surface useful problems. It still leaves you relying on the system you’re checking.

Match verification to the task. Open the cited research. Recalculate the number. Run the code against a case with a known result. Have someone with relevant expertise review an unfamiliar claim.

For example, a market-size figure needs more than a working link. Check that the source measures the same market, geography, and year as the claim.

Advertisement

The useful question is: What could show that this answer is wrong?

4. Record the cleanup as well as the time saved

Choose a recurring task and keep a simple record for the next five attempts:

  • How much correction you expected.
  • What errors or omissions you actually found.
  • How long verification and revision took.
  • Whether the finished work met your requirements.

A draft that appears in 30 seconds may still save substantial time. Counting the work that follows gives you a more honest picture of the benefit—and your ability to predict it.

Repeated surprises are useful information. If the same problem keeps appearing, change the workflow and check whether the change helps.

5. Test your judgment outside your strongest task

Being good at using AI for one kind of work doesn’t establish competence across every use case.

Choose a task where your own knowledge is thinner. Before accepting the result, identify what you can check yourself and what needs another source, a test, or a reviewer.

You might confidently assess an AI-written customer email while needing help evaluating an unfamiliar data analysis. Recognizing that boundary is a skill worth developing.

The goal is to make your confidence specific: I trust this result because these checks passed.

What better AI training should look like

For employers, this suggests a more demanding way to assess progress.

A team can demonstrate frequent use, sophisticated prompts, and attractive outputs without showing how reliably it catches mistakes. Training should give people opportunities to evaluate plausible but flawed work, choose appropriate checks, and compare their expectations with actual results.

Advertisement

An exercise could be as straightforward as reviewing an AI-generated brief with a planted unsupported claim. Does the employee find it? Can they explain why it matters? Do they know how to fix it?

That makes judgment visible.

As organizations delegate more work to AI, they’ll need people who can decide what deserves further review and what is ready to use. The practical opportunity is to build that judgment alongside the ability to generate.

Confidence becomes useful when you can point to what earned it.

Corey Noles

Corey Noles is the Host of The Neuron: AI Explained podcast and Managing Editor of AI and Experimental Content at TechnologyAdvice, where he leads the charge in testing and refining emerging content strategies across the company's portfolio.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.

Stay in the loop

Get notified when we publish new articles.