## Multiple agents

*Agents*

*Last edited: 28 September 2026*

A lead agent splits the task and delegates parts to subagents, each working in its own clean context and returning a short result. It helps with independent parts, such as researching many sources at once, and hurts with shared state, such as a change to one codebase. It usually costs several times more tokens.

**In plain words:** Five researchers will get through ten libraries faster than one, as long as each gets a clear brief and hands back a page of notes. Five programmers fixing the same module without talking to each other will do more damage than one, because each will quietly make different decisions. You pay both teams for every hour.

*Interactive widget on the page: Change the number of subagents and the type of task. Watch the time, the tokens, the lead agent’s context and the quality of the result.*

### Lead agent and subagents

- A subagent is a separate agent loop (see “The agent loop”) that the lead agent calls like a tool: the argument is a brief, the result is a short report. The subagent does not see the lead agent’s history, only its own system prompt and the brief. It may read tens of thousands of tokens of pages or files, and usually hands back 1–2k. This is the context isolation from “Context engineering and memory”.
- The lead agent, also called the orchestrator, plans, hands out the parts, waits and merges the results. In Anthropic’s research system (described in June 2025) it launches subagents in batches and waits for the whole batch. That simplifies coordination, but the slowest subagent sets the pace, and a subagent cannot be corrected while it works.
- Other layouts. Pipeline: fixed stages, where one agent’s output is the next one’s input. Handoff: an agent passes the conversation, with its state, to a specialised agent and drops out, for example customer service handing a case to the returns team. Critic loop: one agent writes, another grades the result against criteria and sends back comments until it passes or the round limit runs out. With a subagent the lead agent keeps control; with a handoff it gives control away.

### When it helps

- Breadth-first, parallel work: the task splits into independent directions, such as many companies, sources or hypotheses. Anthropic reports that Claude Opus 4 as the lead agent with Claude Sonnet 4 subagents outperformed Claude Opus 4 alone by 90.2% on its internal research eval. Parallel subagents and parallel tool calls cut the time of complex queries by up to 90%.
- More tokens per task than one window holds. In the same report, token usage alone explained 80% of the variance in performance on BrowseComp, a benchmark for finding hard-to-locate information. Subagents let you spend those tokens, while only the conclusions reach the lead agent, so its context stays clean.
- Specialisation: each subagent has its own prompt and a narrower set of tools and permissions. From a shorter list the model picks the right tool more accurately (see “Tools (function calling)”). A read-only subagent can read untrusted content without access to risky actions, but its report can still carry an injected instruction to the lead agent (see “Prompt injection”).

### When it hurts

- Cost and latency. In Anthropic’s data an agent uses about 4 times more tokens than a chat, and a multi-agent system about 15 times more. Each subagent pays for its own system prompt, tools and brief. It starts from zero, so it first gathers context the lead agent already had.
- Lost context and duplicated work. A subagent knows only what is in its brief, and every summary loses detail. An early version of Anthropic’s system gave short briefs, such as “research the semiconductor shortage”. One subagent researched the 2021 automotive chip crisis, while two others duplicated each other’s work on 2025 supply chains.
- Conflicting decisions in shared state. Every action carries implicit decisions, and parallel agents cannot see each other’s (Cognition, “Don’t Build Multi-Agents”, June 2025). In their Flappy Bird clone example, one subagent built a Super Mario-style background and another a bird in a different style, leaving the lead agent to glue them together. In code this means the same files edited at once, different names and types, merge conflicts. Anthropic itself notes that most coding tasks have fewer truly parallelisable parts than research.
- Errors compound, and debugging is hard. A bad split by the lead agent breaks every branch at once, and a small change to its prompt can change subagent behaviour unpredictably. Cemri et al. (2025) analysed more than 1,600 traces from 7 frameworks and described 14 failure modes in three groups: system design, inter-agent misalignment and task verification. On popular benchmarks, multi-agent systems often gained very little.

### How to build it

- Start with one agent with tools; OpenAI’s guide gives the same advice. Add more agents once you show, on the same eval set, a gain worth the extra tokens (see “Evals”). For code, a read-only subagent that searches the repository and answers a specific question is usually enough, while a single agent makes the changes. The exception is a large change that splits into independent units, such as a module-by-module migration: each writing subagent works in its own git worktree or sandbox on its own branch and runs the tests, and its change is merged like an ordinary PR (Claude Code’s `/batch` works this way). Each unit still has a single writer, the rule Cognition also arrived at in April 2026.
- The lead agent’s brief is a specification: the goal, output format, tools and sources, boundaries (what not to do, what the others are doing) and an effort budget. Anthropic wrote scaling rules into the prompt: a simple question gets 1 agent and 3–10 tool calls, a comparison 2–4 subagents with 10–15 calls each, complex research more than 10 subagents. For a code change, the shared contract (names, types and interfaces) is agreed before the work is split.
- Pass large results through files or a store, not through messages. The subagent writes an artefact and returns its path with a short description, so the content does not pass through successive summaries. The lead agent also saves its plan outside its context, because on long tasks the window will run out.
- Limits are enforced by code, not by the prompt: the number of subagents, nesting depth, a token and time budget per subagent, the number of critic-loop rounds. Without them, early versions of Anthropic’s system spawned 50 subagents for simple queries. Claude Code (as of September 2026) allows three levels of nesting and at most 20 concurrent subagents by default (see “The coding-agent harness”).
- One trace for the whole task, with a separate span for each subagent: brief, calls, tokens, result. Only then can you tell whether the split, a subagent or the merge failed. Grade the end state, for example the facts and sources in the report or passing tests, not the lead agent’s summary (see “The agent loop”).

### Check yourself

**Question:** When would you build a system with multiple agents, and when would you stick with one?

**Short answer:** Start with a single agent. Add more when the task splits into independent parts, such as researching many sources: a lead agent hands them to subagents, and each works in parallel in its own clean context and returns a short summary. The gain is speed, breadth and a clean lead context. The price is tokens, about 15 times a chat in Anthropic’s data, and coordination. With shared state, such as one change across a codebase, agents make conflicting decisions, so there one agent makes the changes. End-state evals decide.

### Follow-up questions

- **A subagent doesn’t know what the others decided. How do you deal with that?** Either pass it the full context and the decisions made so far, which eats up the gain from isolation, or agree the shared things before the split: names, types, output format, the scope of each part. If that can’t be settled up front, the parts aren’t independent and one agent will do better.
- **The lead agent spawns 50 subagents for a simple question. What do you do?** That was a real bug in an early version of Anthropic’s system. Scaling rules go into the prompt: how many subagents and calls for which type of question. Code adds hard limits on the number of subagents, the depth and the token budget, and evals check the change.
- **How do you pass large results between agents?** Through files or a store: the subagent writes an artefact and returns its path with a short description. Details don’t get lost in successive summaries, and the lead agent’s context stays small.
- **How is a handoff different from calling a subagent?** A subagent works like a tool: the lead agent waits for the result and keeps control. A handoff passes the conversation, with its state, to another agent, which carries on talking to the user itself. A handoff suits routing a case to a specialist; a subagent suits splitting a task into parts.
- **How do you debug and evaluate a multi-agent system?** One trace ID for the whole task and a separate span for each subagent with its brief, calls and tokens. Grade the end state on a fixed set and compare it with a single agent, because a small change to the lead agent’s prompt can change the behaviour of every subagent.

### Sources

- [Anthropic: How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system)
- [Cognition: Don’t Build Multi-Agents](https://cognition.com/blog/dont-build-multi-agents)
- [OpenAI: A practical guide to building agents (PDF)](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf)
- [Cemri et al.: Why Do Multi-Agent LLM Systems Fail? (arXiv, 2025)](https://arxiv.org/abs/2503.13657)
- [Cognition: Multi-Agents: What’s Actually Working (April 2026)](https://cognition.com/blog/multi-agents-working)

Interactive page: https://howaiworks.dev/multi-agent/
