## MCP

*Agents*

*Last edited: 28 September 2026*

MCP (Model Context Protocol) is an open standard for connecting tools and data to model-powered applications. You write an integration once, as an MCP server, and use it in Claude, ChatGPT, an IDE or your own agent. The model knows nothing about MCP: it still sees tool definitions in the prompt and writes calls.

**In plain words:** A wall socket. The kettle maker doesn’t know your wiring, and the electrician doesn’t know the kettle: a shared plug shape is enough. But the socket doesn’t check what you plug into it, so connecting faulty equipment is on you.

*Interactive widget on the page: Connect MCP servers and watch what lands in the model’s context before the user types anything.*

### Why a standard: M + N instead of M × N

- Without a standard, every model-powered application writes its own integration with every system: M applications and N systems mean M × N integrations. With MCP, each system gets one server and each application one client, so M + N is enough. Anthropic published MCP in November 2024 and in December 2025 handed it over to the Agentic AI Foundation under the Linux Foundation, so the standard doesn’t belong to a single vendor.
- MCP standardises the application side: how to discover a server’s tools (`tools/list`) and how to call them (`tools/call`). The model side doesn’t change. The host turns the server’s definitions into ordinary tools in the model API, the model writes a call, and the host forwards it to the right server and hands back the result. The mechanism itself is covered in “Tools (function calling)”.
- MCP connects an agent to tools and data; A2A (Agent2Agent, from Google, a Linux Foundation project since June 2025) connects agents to each other. An A2A agent publishes an Agent Card with its skills, and another agent hands it a task and gets back status updates and artefacts, without seeing its tools or prompt. As of September 2026 A2A has over 150 supporting organisations and integrations in the Google, Microsoft and AWS agent platforms, but no public usage figures. Subagents inside one application don’t need it (see “Multiple agents”); it pays off when the other agent belongs to another team or vendor.

### Host, client, server

- The host is the model-powered application, e.g. Claude Desktop, an IDE or your agent. It creates a separate client for each server, and the host decides what reaches the model. A server doesn’t see the conversation or the other servers; it only gets the calls addressed to it. Messages are JSON-RPC 2.0 over one of two transports: stdio, when the host launches the server as a local process and talks to it over stdin and stdout, or Streamable HTTP for remote servers, where every message is a POST to a single endpoint.
- A server exposes three kinds of things, which differ in who decides to use them. Tools are chosen by the model. Resources, e.g. a file or a database schema, are attached to the context by the application. Prompts, i.e. ready-made templates, are chosen by the user, usually as a slash command. In the other direction, a server can ask the user for input or consent during a call (elicitation): with a form, or, for passwords and payments, with a link opened outside the application. Sampling (the server asks the host’s model for a response), roots (the host points the server to directories) and protocol-level logging are deprecated as of version 2026-07-28. They work for at least 12 more months; the replacements are tool parameters, direct calls to the provider’s API, and stderr or OpenTelemetry. Optional extensions, negotiated through capabilities, add Tasks (a long-running call returns a handle the client polls) and MCP Apps (interactive UI the host renders in the conversation).
- Specification version 2026-07-28, current as of September 2026, made the protocol stateless. There is no `initialize` handshake and no session any more: every request carries the protocol version and the client’s capabilities in `_meta`, and `server/discover`, which every server must implement and a client may call, returns the server’s versions and capabilities. As a result, a remote server scales behind an ordinary load balancer. When the server needs something from the user, it doesn’t send a request of its own; it replies with `input_required`, and the client retries the call with the answer. Older servers (up to version 2025-11-25) start with `initialize` and capability negotiation and keep a session. A client that wants to talk to them detects the version and switches to the old mode. The reverse doesn’t work: a client on 2025-11-25 or earlier has no way to switch forward, so a server that must also serve such clients has to support both eras and keep answering `initialize`.
- A remote server authorises requests with OAuth 2.1 and acts as the resource server. To a request without a token it responds with 401 and the address of its metadata. In the metadata the client finds the authorisation server, takes the user through sign-in and gets a token issued for this one server. The server must not accept a token issued for another service, nor pass the client’s token on to an API (token passthrough). A local stdio server takes its credentials from environment variables.

*Interactive widget on the page: Step through an exchange with a calendar server. Watch which part is the MCP protocol and which is ordinary function calling.*

### Every server costs tokens

- A simple host attaches the definitions of all tools from all connected servers to every request. Anthropic describes five servers with 58 tools that took about 55k tokens before the first question was asked. You pay for this on every turn, and the model makes more mistakes when tools have similar names, e.g. `notification-send-user` and `notification-send-channel`.
- Limit this on the host side: connect only the servers you need and enable only the tools you need from them. With a large catalogue, the host defers the definitions: the model gets a tool for searching tools, and full definitions enter the context on demand. As of September 2026 Claude Code, Codex and Cursor do this by default; Claude Code loads only tool names and server instructions at the start. The MCP docs recommend this mode once definitions take up 1–5% of the window. Don’t reorder or remove tools mid-conversation, because that invalidates the prompt cache (see “Prompt caching”); definitions found by search are appended after the cached prefix, so they don’t break it.
- Design the server around tasks, not API endpoints: one `schedule_event` that finds a free slot and creates the meeting, instead of three wrappers, `list_users`, `list_events` and `create_event` (Anthropic’s example). Return concise results with only the fields the model needs, filtered and paginated on the server side, because a result stays in the context for every later turn. When the host defers definitions, the tool name, description and server instructions decide whether the model finds your tool at all, so put the key facts first. When an agent chains many calls, code mode helps: the model writes a script that calls the tools in a sandbox, and only the final result comes back into the context.

### Someone else’s server, code and text

- A local server runs with your user’s permissions and can read files or SSH keys. A remote one receives every argument the model passes to it. Tool descriptions and tool results go into the prompt as text, and the model can’t tell them apart from your instructions (see “Prompt injection”).
- Tool poisoning is an instruction hidden in a tool description. The model reads it even if it never calls that tool: in every request when the host loads definitions up front, or once a search returns it, while the user usually sees only the name and a shortened description. In April 2025 Invariant Labs showed an `add` tool whose description got an agent in Cursor to read the SSH key and the MCP configuration and send them in an extra parameter. A malicious description can also change how the tools of other, trusted servers are used (shadowing).
- Rug pull: a server changes its definitions after you have approved it. It only takes the next `tools/list` returning a different description, or a new package version doing something other than the previous one.
- A malicious server supplies two parts of the lethal trifecta by itself: untrusted text and an exfiltration channel, because everything the model writes into its tools’ arguments goes to the server’s owner. Private data comes from any other connected server. Overly broad permissions multiply the damage: a token for all repositories turns one successful attack into a leak of everything.
- The defence lives in the host, not in the prompt: an allowlist of servers, only the tools you need from them, pinned versions and fresh approval when definitions change, local servers in a sandbox, tokens with minimal scope widened on demand. Call tools that change or send something only after approval by a human who sees the arguments. Annotations such as `readOnlyHint` are declared by the server itself, so from an untrusted server they mean nothing.

### When you don’t need MCP

- One application with a few internal tools: plain function calling in the application code is enough. MCP adds a separate process or service, a transport, authorisation and versioning, and the M + N gain only appears once several applications use the same integration.
- MCP pays off when an integration has to work in many hosts (your products, your team’s IDEs, your customers’ Claude and ChatGPT) or when you want to use a vendor’s ready-made server instead of writing the integration yourself. The OpenAI API (the `mcp` tool in the Responses API) and the Anthropic API (the MCP connector) connect to a remote server themselves, so you don’t have to write a client.

### Check yourself

**Question:** What is MCP, and when would you use it instead of plain function calling?

**Short answer:** MCP is an open JSON-RPC protocol that standardises the application side of function calling: the host discovers a server’s tools with tools/list and invokes them with tools/call. The model notices nothing; it still sees tool definitions in the prompt and emits calls. The gain is M + N integrations instead of M × N: a server is written once and works in Claude, ChatGPT or an IDE. The cost is the connected servers’ definitions in the context (all of them on every request, unless the host defers them through tool search), and a new trust boundary, because a third-party server is foreign code and foreign text in the prompt. For one app with a few internal tools, plain function calling is enough.

### Follow-up questions

- **You have 30 MCP servers and the agent starts picking the wrong tools. What do you do?** Measure how many tokens the definitions take and which tools overlap. Keep only the tools the task needs, and load the rest through tool search or split them between sub-agents with their own sets. Don’t reorder or remove tools mid-conversation, so as not to break the prompt cache.
- **How do you allow MCP servers in a company without opening a path to data leaks?** An allowlist of approved servers with pinned versions and a diff of the definitions on every update, local servers in a sandbox, and minimally scoped OAuth tokens issued for a specific server. Write and send actions are approved by a human in the host, and every call goes to an audit log.
- **How do tools, resources and prompts differ?** In who decides to use them. The model chooses tools, the application attaches resources, e.g. a file or a database schema as context, and the user picks prompts, usually as a slash command. Not every host supports all three: the MCP connector in the Anthropic API supports only tools.
- **Why did version 2026-07-28 remove sessions and initialize?** So that a remote server scales like an ordinary API. Every request carries the protocol version and the client’s capabilities, so it can hit any instance behind a load balancer with no shared state. The server passes state between calls explicitly: a tool returns a handle, and the model passes it in the next call.
- **You are building an MCP server on top of an existing REST API. How do you approach it?** Don’t map endpoints one to one. Pick a few tasks the agent actually performs and turn them into tools that combine several API calls themselves. Return only the fields needed, filter and paginate on the server side, and send validation errors back as a result with isError and a hint on how to fix the arguments.

### Sources

- [MCP specification, version 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28)
- [MCP: Client Best Practices (tool search, code mode)](https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices)
- [MCP: Security Best Practices](https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices)
- [Invariant Labs: Tool Poisoning Attacks (2025)](https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks)
- [Linux Foundation: A2A after its first year (April 2026)](https://www.linuxfoundation.org/press/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year)

Interactive page: https://howaiworks.dev/mcp/
