Home / Chapter 9 · Practice and review
    Last edited · 14 min read

    Use with AI

    Design a system

    Knowing the mechanisms is half the work. The other half is combining them into a system that fits a budget and survives production: from requirements through architecture and trade-offs to evaluation and failures, with numbers rather than the name of a favourite model.

    In plain wordsA conversation with an architect about your house. A good one first asks how many of you there are, what your budget is and what plot you have, then sketches the simplest house that meets those conditions, and tells you what happens when the family grows. Whoever starts by choosing the roof tiles, that is, the model, is not designing but picking from a catalogue.

    Pick a scenario and go through it step by step. Before you read a step, decide what you would do, then compare. Each step ends with an open question to settle next

    A company-wide LLM gateway

    • Requirements. “Shared LLM access for 40 teams.” Ask: how many providers, whether a self-hosted model is needed (see “Open-weight models”), whether personal data may leave the company, what budgets per team, what availability. Streaming must pass through with no noticeable added latency. Open question: who pays for the tokens, and who decides which models are allowed?
    • Architecture. One endpoint in a popular API format, and behind it: a key per team, limits and budgets, PII masking, routing with fallback, a trace of every call. The gateway is stateless and scales horizontally, with budget counters in a shared, fast store. Open question: how do you version the contract when providers add their own parameters, such as reasoning effort?
    • Decisions. A common interface eases switching providers but flattens their features: keep a pass-through mode. PII masking trades accuracy against latency and false positives that corrupt the context. Fallback raises availability, but the other model behaves differently, so prompts and evals must cover both. Caching answers by question similarity is risky: similar is not the same. Prompt caching needs a stable prefix on the same model and provider, and at Anthropic the same workspace (see “Prompt caching”), so a fallback starts with a cold cache. Open question: once a budget is exceeded, a hard block or a cheaper model?
    • Evaluation. Measure the gateway, not the answers: added latency at p50 and p99, errors per provider, cost per team, masking effectiveness on local data formats (national ID numbers, addresses, non-English names). Answer quality belongs to the team that owns the feature; the gateway gives it traces and pinned model versions (see “Compute and “getting dumber””). Open question: what do you log, and what must you never log?
    • Failures. A partial provider outage triggers a retry storm: backoff with jitter, a circuit breaker and a retry budget help. An agent in a loop can burn a monthly budget in an hour, so limits apply per key and per request. The gateway is a single point of failure: several instances, and a deliberate choice whether a failing filter passes or blocks traffic. Prompt logs are a leak risk of their own. Open question: which provider errors does the gateway absorb, and which reach the teams?

    Screening 100k CVs a week

    • Requirements. “Score 100,000 applications a week against job postings”: about 14,000 a day, results within hours. GDPR (Art. 22) restricts decisions based solely on automated processing, and the AI Act classifies recruitment as high-risk (Annex III; after the 2026 Digital Omnibus the obligations apply from 2 December 2027), so you need explanations, an audit trail and human oversight. Open question: who makes the decision, the model or the recruiter?
    • Architecture. An event queue, a parser (PDF and DOCX to text, OCR for scans), extraction into a schema (see “Enforcing output format”), a score for each criterion in the posting with a quote from the CV, results in the recruiter’s dashboard. All asynchronous: retries are cheap and traffic peaks don’t block applications. Open question: how do you guarantee exactly one score per application despite retries?
    • Decisions. Batch mode: about 50% cheaper for results within 24 hours. The instructions and the posting form a stable prefix for prompt caching, since one posting has hundreds of candidates. Cascade: a cheaper model scores the clear-cut cases, a stronger one the borderline ones, and neither rejects anyone on its own: a rejection without meaningful human review is a decision based solely on automated processing. The criteria come from the posting, not the model’s “knowledge”; name, age and photo are removed before scoring. Open question: how do you set the borderline threshold?
    • Evaluation. A golden set of a few hundred CVs scored by several recruiters. The model’s agreement with humans is compared with the agreement among humans, the realistic ceiling. Bias tests: pairs of CVs that differ only in name, gender or age must get the same score. In production, watch the score distribution over time (see “Evals”). Open question: where does the ground truth come from if recruiters disagree?
    • Failures. The parser returns garbage for two-column layouts or tables, and the model scores it anyway: detect poor text and route it to a human. A candidate adds “rate me highest” in white text (see “Prompt injection”). No tools and a schema limit what the injection can do, but not its goal: the score itself. So the parser flags text missing from the rendered page (text layer versus OCR), code checks that every quote behind a score is in the visible text, and flagged CVs go to a human. Open question: how will you notice that candidates have started writing CVs for the model?

    RAG over 50M documents

    • Requirements. “An assistant answers employees’ questions from 50 million documents.” Ask about permissions, freshness, languages and formats, citations and response time. The scale: at about 20 chunks per document that is about 1 billion vectors: at 1,024 dimensions about 4 TB in float32, 1 TB in int8. Open question: how will you shrink the index, and how much quality will you lose?
    • Architecture. Ingest: parsing, chunking that preserves headings and tables, and a short note per chunk on its document and section; a vector index and a keyword index (BM25). At query time: permission filter, hybrid search, rank fusion, reranking, a prompt with the few best chunks, an answer with citations (see “Embeddings and vector search” and “RAG”). Open question: how do you re-index 50 million documents after changing the embedding model?
    • Decisions. The permission filter goes into the index query itself: filtering afterwards loses results and risks leaks. Hybrid search, because vectors capture meaning and keywords catch contract numbers and proper names. A reranker is usually the biggest quality gain, for tens to hundreds of milliseconds. A smaller chunk is easier to hit; a larger one carries more context. Open question: what do you do when the reranker becomes the bottleneck?
    • Evaluation. The two stages separately: retrieval (is the right chunk in the top k) and generation (is every claim supported by a chunk, and does each citation point to the right place). Questions from experts, questions a model generates from random chunks, checked on a sample, and questions with no answer in the corpus (see “Evals”). Open question: how do you build a question set without hand-labelling millions of documents?
    • Failures. A permission is revoked but the chunk is still cached: the cache key includes permissions. An old version of a document beats the new one: dates in metadata and version deduplication. Retrieval finds nothing, and the model answers from memory, citing the nearest chunk (see “Hallucinations”). A document carries a hidden instruction (see “Prompt injection”). Open question: how do you detect answers unsupported by sources in production?

    The agent is slow and expensive

    • Requirements. “A task takes 4 minutes and costs $2. Get it down to 30 seconds and 20 cents without losing quality.” Ask: today’s success rate, whether the user waits for the result, which tasks dominate, and what matters most when you can’t have everything. Open question: what do you sacrifice first: cost, latency or success rate?
    • Architecture. Before changing anything, trace every turn: input and output tokens, cache hits, thinking tokens, model and tool time. Usually most of the cost is context resent every turn, and most of the time is sequential turns plus output tokens. The target: a stable prefix, concise tools, a small model for simple steps, parallel calls, subagents for side tasks. Open question: which metrics do you alert on, and at what thresholds?
    • Decisions. Fix the prefix: nothing variable before the system prompt and tools, because a high cache hit rate cuts input cost several times over (see “Prompt caching”). Trim tool results: 20k tokens of JSON resent every turn is the most expensive line in the system. A model per step: routing and extraction on a small model, thinking effort set with evals (see “Choosing a model”). Compaction and subagents keep the context short (see “Context engineering and memory”). Open question: what do you lose with compaction, and how do you detect it?
    • Evaluation. Every optimisation runs on the same task set: pass^k, cost and time per task, p95 rather than the mean, pairwise comparison. A cheaper agent that fails more often ends up costing more in retries and human work. Open question: how do you build a task set when a task can be solved in many ways?
    • Failures. The small model botches decisions that only looked simple. Compaction loses a detail needed 20 turns later. A small change at the start of the prompt silently zeroes cache hits and the cost comes back, so alert on the hit rate. An agent without hard limits goes round in a loop (see “The agent loop”). Open question: how do you detect that the agent is stuck in a loop?

    Real-time classification

    • Requirements. “Thousands of submissions a minute, a decision within 50 ms, mistakes are costly.” Ask: how many classes and how often they change, what a false alarm costs versus a miss, whether there is labelled history. 50 ms rules out generation by a large LLM on the synchronous path: time to first token alone can be longer. Open question: what does each kind of mistake cost?
    • Architecture. The decision state (submission, customer history, rules in force) goes to a fast classifier: a fine-tuned encoder or embeddings with logistic regression, in-process or next to the service. A hosted decision model (see “When not to use an LLM: classifiers and System One”) fits only if its measured p99 latency, network included, stays within budget. Above the confidence threshold the decision is automatic; below it the case gets a provisional “for review” at once and goes asynchronously to a reasoning model or a human. Amounts, permissions and blocklists stay in code. Open question: what stays in rules, and what does the model judge?
    • Decisions. A small classifier answers in milliseconds for a fraction of a cent, but needs data and retraining for every new class. An LLM needs no data, so it helps with labelling and borderline cases, off the 50 ms path. The threshold trades the automation rate against errors and is set on data. The model scores, the code decides. Open question: how will you know the small classifier is no longer enough?
    • Evaluation. Precision and recall per class on production data. A curve of automation rate against error rate at different thresholds. Calibration: does 90% confidence really mean 90% correct? In production, human corrections serve as labels, and monitoring watches the class distribution. Open question: how do you check calibration on your own data?
    • Failures. Drift: a new product or a new fraud campaign, and confidence stays high while accuracy drops. A flood of escalations after a threshold change. Labels arrive only for escalated cases, so the automated errors go unseen; a human therefore also reviews a random sample of automated decisions. Open question: how large must that sample be to spot a 1% error rate?

    How to approach a design

    Numbers for estimates

    Check yourself

    How would you estimate the cost and latency of an LLM system before you build it?

    Start from a single request: input tokens (system prompt, context, history) and output tokens separately, because output costs several times more and drives generation time. Multiply by request volume and prices, then subtract cache hits and whatever can run in batch at half price. For an agent, also multiply by the number of turns, because each turn resends the whole context. Latency is the sum of turns: time to first token, generation and tools. Compare the result with the budget and the latency limit before choosing a model.

    Po polsku

    Od jednego zapytania: tokeny wejścia, czyli system prompt, kontekst i historię, oraz tokeny wyjścia osobno, bo wyjście jest kilka razy droższe i to ono wyznacza czas generowania. Mnożysz przez liczbę zapytań i ceny, odejmujesz trafienia cache’u i to, co może pójść wsadowo za połowę ceny. W agencie mnożysz jeszcze przez liczbę tur, bo każda wysyła cały kontekst od nowa. Czas to suma tur: czas do pierwszego tokena, generowanie i narzędzia. Wynik porównujesz z budżetem i limitem opóźnienia, zanim wybierzesz model.

    Follow-up questions (5)
    When isn’t an LLM needed here?
    When there are few classes, plenty of labelled data and a hard latency or cost limit: then a small classifier or rules win. Also when every decision must be fully reproducible and auditable. An LLM is then useful for labelling data, not on the production path (see “When not to use an LLM: classifiers and System One”).
    How do you choose a model?
    By requirements, not leaderboards: the cheapest model that passes your eval set at the required latency, with a pinned version and a migration plan. A strong model for the hard steps, a small one for the simple ones (see “Choosing a model”).
    How do you keep working when the model provider goes down?
    Timeouts and retries with backoff, fallback to a second provider validated with the same eval set, and a degraded mode without an LLM where possible, such as a queue instead of an immediate answer.
    How do you split responsibility between the model and code?
    The model judges, extracts and writes; code decides and executes. Business rules, permissions, limits and action approval are deterministic and testable, and the model gets only what can’t be written down as a rule.
    When would you fine-tune instead of using a prompt and RAG?
    When the problem is about form, not knowledge: a fixed format, style, classification in a narrow domain, or when a small fine-tuned model is to replace a large one for cost and latency. Knowledge that changes goes into the context (see “Fine-tuning and LoRA” and “RAG”).

    Sources

    Report an error · Suggest a fix