Knowledge · AI at work · 8 minutes

AI hallucinations: what they are and can you trust ChatGPT?

AI hallucinations are confident but made-up answers. Real examples, why they happen, how to check chatbot answers at work and whether AI text detectors work.

Last updated: October 6, 2026

What is an AI hallucination?

An AI hallucination is content generated by an AI model that is stated with confidence but is false or does not match the question. Technical documents also call it "confabulation".

The US National Institute of Standards and Technology, in its Generative AI Profile (NIST AI 600-1, July 2024), defines confabulation as the production of confidently stated but erroneous or false content that may mislead users. NIST also counts answers that contradict the prompt or something the model said earlier in the same conversation.

OWASP, the application security community, lists the problem among its top ten risks for applications built on large language models as LLM09:2025 Misinformation. It points to both sides of the issue: the model makes things up, and people trust it too much.

AI hallucination examples

In office work the most common case is invented sources. The chatbot gives an article title, author and year, and the publication does not exist. Laws work the same way: the article number is right and the content is not, or the other way round.

OWASP describes several known cases. Air Canada's chatbot gave a customer wrong information about its refund rules, and a tribunal held the airline responsible for what its chatbot said. Lawyers in the United States filed a brief citing court decisions that ChatGPT had invented. Coding assistants suggest libraries that do not exist, and attackers register packages under those names and wait for someone to install them. More on that in our piece on vibe coding security.

A well-known business example is a Deloitte report for the Australian Department of Employment and Workplace Relations worth 440,000 Australian dollars. According to the Australian Financial Review of 5 October 2025, it contained nonexistent academic references and a made-up quote from a Federal Court judgment. The corrected version removed more than a dozen false footnotes and disclosed the use of Azure OpenAI GPT-4o, and the firm agreed to repay the final instalment of its fee.

Why does AI hallucinate?

A language model does not look facts up in a database. It predicts the next word based on patterns in its training data. NIST says plainly that confabulations are a natural result of this design: the same statistical prediction often gives correct answers and sometimes false ones.

Researchers from OpenAI and Georgia Tech add a second reason in "Why Language Models Hallucinate" (September 2025). Training and evaluation reward guessing more than admitting uncertainty, like an exam where a blank answer scores zero. So the model learns to always answer, even when it does not know.

Hallucinations are most frequent where the model saw little data: narrow specialisms, local regulations, little-known companies and people, recent events, exact numbers and dates.

Can you trust ChatGPT and other chatbots?

It depends on the task. A chatbot works well where the result is easy to judge: improving style, drafting an email, summarising a text you pasted yourself, brainstorming. It works less well where you have to take the answer on faith: laws, tax rates, medical facts, quotes, sources.

Does ChatGPT lie? Not in the human sense, because it has no intention to mislead you. But the effect can be the same: a confident tone and no signal that it does not know. So confidence in an answer is not evidence that it is true.

A simple rule helps: the harder the result is to check and the bigger the cost of a mistake, the less trust. You can fix a draft post on the spot. An answer about an official deadline, a drug dose or contract terms has to go through a source or a specialist. Web search built into a chatbot lowers the risk because the answer rests on real pages, but it does not remove it: the model can still misread a page or merge two sources into one false claim.

The safety of the data you type into a chatbot is a separate question. We cover it in is ChatGPT safe.

How to check AI answers at work and reduce hallucinations

1. Check every law, quote, number, date and name at the source. Laws in the official journal or EUR-Lex, companies in the business register, publications in the publisher's database. If you cannot find the source, treat the information as made up.

2. Open the links the chatbot gives you. Sometimes the link works but the page says something different from the answer. If the chatbot pointed you to an unknown website, you can check the domain for free before visiting it.

3. Give the model material instead of asking it to answer from memory. Paste a document (if you are allowed to) and ask for an answer based only on it. This is a simple form of retrieval-augmented generation, which OWASP lists first among ways to limit hallucinations.

4. Let the model say it does not know. Add "if you are not sure, say you do not know" to your prompt and ask it to mark which parts are certain and which are guesses.

5. Decide who in your company checks AI output before it goes outside. A person is responsible, not the tool. This rule belongs in your company AI use policy.

6. With code, check that a suggested library exists and who publishes it before you install it.

How to tell if a text was written by AI

In short: you cannot tell for sure. AI text detectors give a probability, not proof, and they are often wrong.

The Stanford study "GPT detectors are biased against non-native English writers" (Liang et al., 2023) tested seven popular detectors. Essays by non-native English writers were flagged as AI-generated in 61% of cases on average. The same detectors were easy to fool by simply rewording the text. The authors warn against using them to judge people.

Technical marking is a more reliable route. From 2 August 2026, Article 50 of the EU AI Act requires providers of systems that generate text, images, audio and video to mark the output in a machine-readable format. Systems placed on the market earlier have until 2 December 2026. Such marking disappears once a text is rewritten, so it does not settle the question either.

If you suspect a document was written with AI, do not guess the author. Check the content: do the sources exist, do the numbers add up, can the author explain the text. That assesses quality, which is usually what matters.

In short

  • AI hallucinations are confidently stated false content. Every chatbot produces them because of how language models work.
  • They most often affect sources, laws, quotes, numbers and little-known topics.
  • Trust a chatbot with tasks whose results are easy to judge. Check facts at the source.
  • AI text detectors are not reliable and are especially wrong about texts by people writing in a second language.
  • In a company, decide who checks AI output before it goes outside.

Have a website, app or email address that looks suspicious?

Frequently asked questions

Does ChatGPT lie?

Not on purpose, but it can state false information with full confidence. The model predicts likely words rather than checking facts, so laws, quotes and sources need to be verified.

What is another word for AI hallucination?

NIST documents use "confabulation", and OWASP covers the problem under "misinformation". Both describe the same thing: confident but false output.

How can I reduce AI hallucinations?

Give the model source material and ask it to answer only from that, let it admit it does not know, and check facts at the source. You cannot eliminate hallucinations completely.

How do you make an AI hallucinate?

The easiest way is to ask about things the model cannot know: a book that does not exist, a narrow local regulation, an exact figure from an obscure report. It is a useful test before rolling out a tool in a company, because you see whether it admits uncertainty.

What happened with Deloitte and AI hallucinations?

A 440,000 Australian dollar Deloitte report for an Australian government department contained nonexistent sources and a made-up court quote. According to the Australian Financial Review, the firm corrected the report, disclosed its use of GPT-4o and agreed to repay the final instalment of its fee.

Are AI text detectors reliable?

Not reliable enough to judge anyone. A 2023 Stanford study found that seven detectors flagged essays by non-native English writers as AI-generated in 61% of cases on average.

Sources

  1. NIST AI 600-1: Artificial Intelligence Risk Management Framework, Generative AI Profile (July 2024)
  2. OWASP: LLM09:2025 Misinformation
  3. Kalai, Nachum, Vempala, Zhang: Why Language Models Hallucinate (arXiv, 4 September 2025)
  4. Liang et al.: GPT detectors are biased against non-native English writers (arXiv, 2023)
  5. Australian Financial Review: Deloitte to refund government after admitting AI errors in $440k report (5 October 2025)
  6. Regulation (EU) 2024/1689 (AI Act), Article 50
  7. Regulation (EU) 2026/1744 amending the AI Act (Digital Omnibus)

Accurate as of the article's last update. Laws and vendor terms change, so check the source before you decide.

See also