Home / Chapter 6 · Quality and security
    Last edited · 8 min read

    Use with AI

    Evals

    An LLM’s output is random and sensitive to small prompt changes, so “I checked it on three examples” means nothing. An eval is a fixed set of cases with automated grading, run after every change to the prompt, model and tools, plus quality measurement on production traffic.

    In plain wordsTests in CI for code that answers slightly differently every time. Instead of “pass or fail” you count how many times out of five it passed, and some assertions can’t be written as string comparisons, so a second model checks them, once you have checked that model.

    A revised prompt v2 raises the score from 5 to 7 out of 8. Check in the table what that number hides, then run each case five times and see which change was noise

    passed, prompt v1
    passed, prompt v2
    cases broken by v2

    Illustrative data: a classifier routing support tickets to teams, eight cases. Red rows are cases v2 breaks; greyed-out rows show no meaningful change.

    The set and the grading

    Noise and repeated runs

    Offline, CI and production

    Check yourself

    How do you check that an LLM system works and that a change hasn’t broken it?

    Build a fixed set of cases from real failures: read traces, name the error types and turn each into a test. Grade with the cheapest method that suffices: exact match, code checks, an LLM judge validated against expert labels, humans. Run every case several times because outputs are random, and compare versions pairwise. The regression set gates every prompt, model and tool change in CI. In production, grade a sample of traffic with a reference-free judge and watch user signals, and feed failures back into the set.

    Po polsku

    Zbuduj stały zestaw przypadków z prawdziwych błędów: czytaj trace’y, nazwij typy pomyłek i każdy zamień w test. Oceniaj najtańszą wystarczającą metodą: dopasowanie, test w kodzie, sędzia-LLM sprawdzony na ocenach eksperta, człowiek. Każdy przypadek puszczaj kilka razy, bo wynik jest losowy, a wersje porównuj w parach. Zestaw regresji blokuje w CI każdą zmianę promptu, modelu i narzędzi. Na produkcji oceniaj próbkę ruchu sędzią bez wzorca i patrz na sygnały użytkowników, a nieudane przypadki wracają do zestawu.

    Follow-up questions (5)
    How do you start without data?
    With error analysis on 50–100 real examples: read the outputs, name the error types and turn them into test cases. 20–50 cases are enough to start, as long as they come from real failures.
    When can you trust a model as a judge?
    When its grades agree with an expert’s on a held-out sample, measured separately for good and bad answers. A judge that passes everything also scores high accuracy when most answers are good. Give it one criterion and a yes/no verdict, and when comparing two answers, swap their order.
    How do you set up a CI gate that doesn’t block on noise?
    Keep stable regression cases that score close to 100% in the gate, and run each several times. Block when a critical case clearly drops or the mean drops by more than the noise margin. Move cases that flicker without any change into quarantine and investigate them instead of loosening the threshold.
    How do you evaluate an agent?
    With a test of the end state, not the agent’s summary, running each task several times, with pass^k as the reliability metric. Add cost, number of turns and time per task. Traces show where the agent goes astray. You don’t grade the path rigidly, because several routes lead to the goal, but you do check it for forbidden calls and policy violations.
    You’re switching to a cheaper model. How do you decide?
    The same set on both models, pairwise and with repeats, with a separate score for each case type, plus cost and latency. Then a canary or shadow run on part of the traffic and a comparison of production signals. The prompt often needs retuning for the new model, so you compare the best version of each.

    Sources

    Report an error · Suggest a fix