Home / Chapter 8 · Choosing a model
    Last edited · 6 min read

    Use with AI

    When not to use an LLM: classifiers and System One

    When a system needs a decision rather than text (a label, a score, yes or no), an LLM still writes it token by token, and its probabilities are often missing or poorly calibrated. A non-generative model is often faster, cheaper and better calibrated; an LLM wins when the answer space is open or the task needs reasoning. TypeSafe’s System One serves here as the worked example of a new model type.

    In plain wordsThe difference between writing half a page of justification and ticking boxes on a form. A form needs no writer, but it only has the boxes someone thought of in advance.

    Run both lanes: the LLM writes a three-field JSON token by token, Jev (TypeSafe’s System One model) gets the same three questions in a single call

    State: the ticket “You charged me twice, I want a refund!”. Questions: which team, is it a refund request, how frustrated is the customer.

    LLM

    writes JSON, one model pass per token

    Passes: 0

    Jev

    state and three questions in one call

    Calls: 0

    An illustrative animation, not a benchmark. Each LLM step is one decode pass after the prompt has been processed; a model that reasons before the JSON writes hundreds of tokens more. The Jev side shows one API call; the vendor does not describe how much work the model does inside it.

    The options for a decision

    Worked example: System One (vendor claims)

    Move the threshold: decisions above it run automatically, those below go to an LLM or a human. Watch how many mistakes the automation lets through

    automatedescalated
    decisions made automatically
    escalated to an LLM or a human
    expected mistakes among automated decisions, if the probabilities are calibrated

    The dots are the probability of the chosen answer for 12 illustrative tickets. With good calibration, decisions at probability 0.9 are wrong 10% of the time on average. That does not tell you whether this particular decision is right.

    Fitting it into a system

    Check yourself

    When does a classifier or a decision model such as System One beat an LLM, and when doesn’t it?

    A non-generative model wins when the answer is a label, a score or yes/no from a closed list and volume, latency or calibrated probabilities matter. A fine-tuned encoder answers in milliseconds but needs labelled data for each task; zero-shot NLI encoders need none. An LLM needs no data but writes the decision token by token, and its logprobs are often unavailable or poorly calibrated after post-training. Decision models such as TypeSafe’s Jev claim typed answers with calibrated probabilities in one call, which has to be measured on your own data. An LLM wins for text, reasoning, tools and open answer spaces. The pattern: the fast model settles confident cases and escalates the rest to an LLM or a human.

    Po polsku

    Model niegeneratywny wygrywa, gdy odpowiedzią jest etykieta, ocena albo tak/nie z zamkniętej listy, a liczą się wolumen, opóźnienie albo skalibrowane prawdopodobieństwa. Dostrojony enkoder odpowiada w milisekundach, ale wymaga oznaczonych danych dla każdego zadania; enkodery zero-shot oparte na NLI ich nie potrzebują. LLM obywa się bez danych, ale pisze decyzję token po tokenie, a jego logprobs często są niedostępne albo po post-trainingu źle skalibrowane. Modele decyzyjne, takie jak Jev od TypeSafe, obiecują typowane odpowiedzi ze skalibrowanymi prawdopodobieństwami w jednym wywołaniu, co trzeba zmierzyć na własnych danych. LLM wygrywa przy tekście, rozumowaniu, narzędziach i otwartej przestrzeni odpowiedzi. Wzorzec: szybki model rozstrzyga pewne przypadki, a resztę eskaluje do LLM-a albo człowieka.

    Follow-up questions (4)
    When is an LLM still the better choice?
    When you need text, code, an explanation of the decision, multi-step reasoning, arithmetic or tool calls. Also when the space of answers cannot be closed into a list of options up front, or the volume is too low to justify collecting labelled data.
    How do you check calibration?
    On your own labelled data: split the decisions into bins by stated probability and in each bin compare it with the actual accuracy. Report the weighted average gap as ECE. Do it separately for each question type and language.
    How does a decision model differ from an LLM classifier with logprobs?
    A classifier that emits a single label token is one decode step after prefill, so it is fast too. But each question is a separate request or a longer output, logprobs are often unavailable (never in the Claude API; in OpenAI’s current models only with reasoning off), and post-training does not optimise them for calibration and usually spoils what pretraining gave. Measuring accuracy, calibration, latency and cost on the same dataset settles it.
    How would you set the escalation threshold?
    From costs: automate when (1 − p) × cost of a mistake is lower than the cost of escalation. Then check on data that the actual error rate at that threshold matches the expected one, and repeat that after every model version change.

    Sources

    Report an error · Suggest a fix