Your AI assistant is lying to you.

Robert Cizmas
September 2, 2026
Blog

Your AI assistant has a reassuring habit. Ask it a question and it gives you an answer. Ask for a source and it produces one. Ask whether it is sure and, quite often, it finds a fresh way to sound even more certain.

That certainty is useful when you need to get past a blank page. It is less useful when the answer contains a number pulled from the wrong row, a quote that never appeared in the source or a neat explanation for something the model cannot actually know.

Calling this “lying” is technically unfair. ChatGPT does not have a secret agenda, and your Copilot is not plotting against you between meetings. A language model generates a likely response from the information and patterns available to it. It can produce a false answer without understanding that it is false.

For the person reading it, the distinction offers limited comfort. You still have a confident answer on your screen. You still have to decide whether to use it.

Why AI would rather guess than say “I don’t know”

OpenAI published a paper in September 2025 examining why hallucinations persist even as models improve. Its researchers argued that common training and evaluation methods reward guessing. If a model leaves a question unanswered, it gets no point. If it guesses, it might get lucky. Over time, that incentive encourages plausible answers where an honest expression of uncertainty would be more useful.

OpenAI is explicit that ChatGPT still hallucinates. Newer models do it less often, particularly when reasoning, but the problem has not disappeared.

This becomes harder to manage because bad AI output rarely arrives wearing a warning label. The grammar is clean. The structure makes sense. The source may even be real while the claim attached to it is wrong.

In a 2025 BBC study, journalists assessed answers from ChatGPT, Copilot, Gemini and Perplexity to questions about news. They found significant issues in 51% of the answers. Among responses citing BBC material, 19% introduced factual errors, while 13% of the quotes were altered or absent from the cited article.

That study concerned news, so it cannot tell us how often your quarterly analysis, campaign plan or client proposal contains an error. It does show why a citation cannot finish the checking process. A valid link and a valid conclusion are different things.

The more convincing AI becomes, the more work your judgement has to do

The obvious response is to be careful. The difficulty is that people are not especially good at noticing when their own level of care has dropped.

Microsoft Research and Carnegie Mellon University surveyed 319 knowledge workers for a CHI 2025 study, collecting 936 examples of AI use at work. Higher confidence in generative AI was associated with less critical thinking. Higher confidence in the worker’s own ability was associated with more of it. The work of thinking also changed: people spent more effort verifying information, integrating responses and retaining responsibility for the task.

In other words, AI can remove part of the production work while increasing the importance of review. If you save ten minutes drafting and spend fifteen minutes repairing a polished mistake, the cheerful little sparkle beside the text box will not mention it.

A study by METR produced an even stranger result. Sixteen experienced open-source developers worked on 246 real issues in codebases they already knew, sometimes with AI tools and sometimes without them. Before the study, they expected AI to make them 24% faster. Afterwards, they still believed it had made them 20% faster. The measured result was a 19% slowdown when AI was allowed.

METR is careful about the limits of that finding. It does not prove AI slows down every developer, let alone every profession. The tools have also continued to improve. What the study demonstrates is the gap that can exist between feeling more productive and being more productive.

AI is pleasant to use. It responds immediately, removes the discomfort of starting and turns messy thoughts into respectable prose. Those benefits are real. They can also make capability difficult to judge from experience alone.

Using AI every day is not the same as using it well

Most people assess their AI ability through frequency. If you use ChatGPT every day, know how to write a long prompt and have tried a few different models, you probably feel reasonably capable.

Work asks more of you. Can you give the model the context it needs without burying the task? Can you recognise an unsupported claim? Can you check a calculation against the underlying data? Do you notice when the answer has drifted away from the original goal? Most importantly, can you decide when AI is the right tool and when doing the work yourself would be faster or safer?

These skills are easy to agree with and harder to demonstrate. Nearly everyone says they check important outputs. Fewer people can spot a subtle error inside a plausible answer when there is no helpful red flag beside it.

That is why we built the AQ Assessment.

AQ measures how you work with AI through behavioural questions and practical scenarios, including AI-generated outputs with problems for you to identify. It adapts to your role because good AI judgement looks different in finance, marketing, product management and data analysis. The assessment takes around four minutes and gives you a score from 1 to 10, followed by a clearer view of what you already do well and what you should improve next.

It is not a test of how many AI terms you know. It is a check on whether your current habits help you produce work that deserves your confidence.

Your AI assistant will continue to give you answers. The useful question is whether you know what to do with them.

Test your AQ for free. It takes around four minutes.