Home / Chapter 6 · Quality and security
    Last edited · 7 min read

    Use with AI

    Hallucinations

    A model has no “I don’t know” mode: it always writes a plausible continuation, in the same confident tone whether it knows the answer or is guessing. Hallucinations cannot be switched off; they can be reduced and measured.

    In plain wordsA student in an oral exam who never says “I don’t know”. When they know the answer, they speak fluently and confidently. When they don’t, they speak just as fluently and confidently, so you can’t tell the difference from the tone.

    Turn on a source in the context and permission to say “I don’t know”, separately and together. Watch which fabrications disappear, which remain and how many answers you lose along the way

    correct
    fabricated
    “I don’t know”
    errors among the answers given

    An illustrative simulation, not a measurement of any model: the answers were written by hand to show typical behaviour. Wrenfold, its museum, the parish chronicle and the paper by Lindqvist and Brandt are invented. The set is deliberately made up mostly of hard questions, so the proportions say nothing about how well real models do. Rows whose verdict changed after a click are highlighted.

    Where they come from

    Types

    What helps, and its limits

    Confidence threshold and measurement

    Move the confidence threshold below which the model says “I don’t know”. Compare the score under 0/1 grading and under grading with a penalty for errors, and find the point where abstaining pays off

    of questions answered
    errors among the answers
    points per 100 questions under 0/1 grading, as in most benchmarks
    points per 100 questions when an error costs −1

    An illustrative model: 1,000 questions, half easy and half hard, and perfect calibration, meaning that at 70% confidence it is right 70% of the time. Illustrative numbers, not a benchmark result. Grading with a penalty for errors is the proposal from Kalai et al., “Why Language Models Hallucinate” (2025).

    Check yourself

    Where do hallucinations come from, and how do you reduce and measure them in production?

    The model always produces a plausible continuation and has no built-in “I don’t know”, so on rare facts it guesses in the same confident tone. Training and benchmarks graded 0/1 reward a hit and give zero for abstaining, so guessing pays. It can’t be switched off, only reduced: sourced context with citations, permission to say “I don’t know”, tools for facts and arithmetic, and checking citations in code. Lower temperature just repeats the same guess. Measure errors, abstentions and faithfulness to sources separately, and set the answer threshold by the cost of a mistake.

    Po polsku

    Model zawsze generuje prawdopodobny ciąg dalszy i nie ma wbudowanego „nie wiem”, więc przy rzadkich faktach zgaduje tym samym pewnym tonem. Trening i benchmarki oceniane 0/1 nagradzają trafienie, a za wstrzymanie dają zero, więc zgadywanie się opłaca. Nie da się tego wyłączyć, można ograniczać: źródła z przypisami w kontekście, zgoda na „nie wiem”, narzędzia do faktów i rachunków, sprawdzanie cytatów kodem. Niższa temperatura tylko powtarza ten sam strzał. Mierz osobno błędy, wstrzymania i wierność źródłom, a próg odpowiadania ustawiaj pod koszt pomyłki.

    Follow-up questions (4)
    Can you detect a hallucination from logprobs?
    Partly. Low confidence can be a signal, but a model can be confident in an error, confidence is spread across different phrasings of the same answer, and post-training hurts calibration. A better signal comes from sampling several answers and checking whether they mean the same thing.
    Why does a RAG system still make things up?
    Retrieval always returns something. If the chunk doesn’t contain the answer, the model often fills the gap from its weights and attaches a citation to the nearest source. On top of that, it can distort the chunk right in front of it.
    The agent reports “tests pass”, but they don’t. What do you do?
    Trust the state, not the report. Code runs the tests and sends the result back to the model, and a tool, not the model’s claim, checks the task’s completion condition. In evals you grade the end state, and runs with a false report go into the set as new cases.
    How do you measure hallucinations in your own system?
    A set of questions with verified answers plus questions whose answer isn’t in the sources. Count correct, wrong and abstained answers separately, and for RAG also faithfulness: whether every claim is supported by the supplied chunk.

    Sources

    Report an error · Suggest a fix