LetterMCP

Multi-Agent System Design Principles

Identity tracing through agent chains is MCP's missing foundation for production security.

Reporter · · 10 min read · Updated
Cover illustration for “Multi-Agent System Design Principles”
Agentic AI Foundations · August 25, 2026 · 10 min read · 2,283 words

Multi-agent systems fail in production when you treat security as something to add after the agents are already running. An AI agent doesn't just generate text, it acts. It reads files, writes to databases, sends emails, executes code, and calls other agents on your behalf. A text-generation failure produces a bad sentence. An action-capability failure produces a deleted database, a leaked credential, or a transaction nobody approved. That distinction is why agentic security can't be patched in later the way a web app vulnerability can. Anthropic introduced the Model Context Protocol in late 2024 to give enterprises a standard way to connect agents to tools, databases, file systems, and APIs, and adoption outran the governance needed to control it.

In a normal application, one login covers one request, and authorization has one clean answer: did this user have permission to do this thing? Multi-agent chains break that model. Each agent in the chain is making decisions on behalf of the one before it, and the one before that, back to whatever person or process kicked things off. If there's no way to trace identity through every hop, nobody can answer the basic question an auditor or a security team needs answered: who authorized this action, through what path, with what scope. The Coalition for Secure AI's analysis following a major industry security conference put it bluntly, calling identity "the load-bearing wall of MCP security," and found that most production deployments haven't poured the foundation that wall needs to stand on, a gap already present in a lot of systems running in enterprise environments right now.

The seven attack surfaces MCP opens in multi-agent chains

Diagram: MCP's Seven Attack Surfaces. Visualizes: Show a ranked or sequential list of the seven attack surfaces MCP opens in multi-agent chains, ordered roughly by severity as described in the article: (1) Tool Poisoning — malicious instructions…

MCP's design opens attack surfaces that don't map cleanly onto older threat models, and each one gets more dangerous as agents start chaining calls to other agents.

The most severe confirmed vector is tool poisoning. A malicious MCP server hides instructions inside tool descriptions, and because those descriptions land directly in the model's context window, the model treats them as trusted input rather than untrusted data. Invariant Labs built a full proof of concept against Cursor and pulled SSH private keys off a machine without the user ever knowing, because the tool description shown in the interface was truncated and the malicious text stayed hidden. In April 2026, OX Security disclosed a systemic flaw sitting at the core of Anthropic's own MCP implementations, spanning the Python, TypeScript, Java, and Rust SDKs, a problem that rippled through a supply chain counting hundreds of millions of downloads and left a large number of deployments exposed. Researchers running benchmarks against real-world MCP servers found that tool poisoning succeeded in the majority of attempts across the major LLM agents tested.

Rug pull attacks exploit a different gap: the one between the moment a user approves a tool and every moment after. A tool can look benign at approval time, then have its description or behavior altered later, with no new approval prompt triggered. The user's one-time review never gets revisited, so trust granted once stays in effect indefinitely, even after the tool itself changes.

Prompt injection becomes a chain reaction in multi-agent workflows. One agent's output feeds the next agent's input, so malicious content injected anywhere in the chain can propagate unauthorized actions, data exfiltration, or hijacked control flow through every downstream process. The EchoLeak attack against Microsoft 365 Copilot, tracked as CVE-2025-32711, showed what this looks like in practice: zero-click exfiltration of emails and internal files that Copilot could access, triggered through indirect prompt injection via its retrieval context.

An MCP server acting as an OAuth proxy creates a confused deputy when it fails to check authorization context on each request. The attacker doesn't need to steal anyone's credentials directly, they just need to trick the intermediary into using someone else's credentials on their behalf. CoSAI's research points to the Asana incident in May 2025 as a textbook case: a tenant isolation flaw let data cross organizational boundaries, contaminating a large number of enterprise accounts with data that belonged to other tenants.

Cross-server shadowing adds a more mundane but equally dangerous risk. Attackers register servers with names deliberately close to legitimate ones and bank on typosquatting and dependency confusion to get the model to call the wrong tool.

Supply chain exposure compounds all of this, because MCP ecosystems run on open-source packages and connectors that may never get audited. A scan by Knostic found 1,862 exposed MCP servers on the open internet, and of the 119 it sampled closely, every single one handed over a full tool listing to a request that was never authenticated.

Researchers have also shown that MCP can work as command-and-control infrastructure for offensive agent swarms, and its traffic looks legitimate enough to slide past standard detection tools. And the raw volume of known vulnerabilities is climbing fast: researchers disclosed dozens of CVEs against MCP implementations across the Python, TypeScript, Java, and Rust SDKs in early 2026, and Microsoft had to patch an Important-rated flaw, CVSS 8.8, in its own MCP servers as part of its March 2026 security release.

MCP's authentication gap as the structural root cause

MCP shipped in November 2024 with no authentication framework built in at all. Every one of the attack surfaces described above exists, in part, because that gap was there from day one. That's an architectural fact.

The protocol specification has moved to close the gap. Under the 2026-07-28 spec, OAuth 2.1 is now mandated, using MUST language, for remote HTTP-based MCP servers. That's a real improvement, and it's worth taking seriously as a sign the ecosystem is maturing. But a spec update changes what a correct implementation looks like going forward. Servers already running without it don't get fixed retroactively, and enforcement still depends on what each individual deployment actually does.

The NSA's cybersecurity information sheet on MCP, published in May 2026, described the protocol's design as "flexible and underspecified," comparing it to the early days of web protocols before security conventions caught up to usage. Part of what makes MCP unusual is that it reverses a pattern most security teams are used to: servers can query and execute actions for clients, rather than the other way around, and that reversal opens attack paths that remain largely untraced.

Dynamic Client Registration, defined in RFC 7591, matters because AI agents need to register with authorization servers while they run, not ahead of time through a manual process. Without it, every client in use today, Claude, Cursor, Windsurf, ChatGPT, Copilot, and the long list of others, would need to be pre-registered by hand. Most legacy identity providers were never built to support that kind of runtime registration, which is exactly the capability agentic systems depend on.

The protocol stack taking shape for 2026 reflects this shift: OAuth 2.1 with PKCE for anything browser-based, JWT bearer assertions for service-to-service calls, MCP itself for tool invocation, and signed agent identity tokens that carry the full delegation chain from end to end.

Some teams push back with a reasonable-sounding argument: bolt authentication on after deployment, once things are running and the priorities are clearer. The problem with that approach is structural. Authentication added after the fact gets added agent by agent, server by server, team by team. It doesn't compose across a multi-agent system. It doesn't produce one unified audit trail. And it does nothing to stop an agent from tricking another into misusing its credentials, since that vulnerability is baked into how the system is architected.

Treating every agent as a first-class identity with a bounded, non-transitive permission scope

Diagram: Permission Scope Can Only Shrink Down the Chain. Visualizes: Illustrate the delegation rule: when Agent A calls Agent B, which calls Agent C, each hop must receive a narrower permission scope than the one before — never equal, never broader.

The first concrete design principle follows directly from the auth gap: every agent needs its own unique, cryptographically verifiable identity, and its permissions need to stay scoped to its specific task rather than growing as it passes work along to other agents.

Standard SaaS identity and access management was built for human users who log in with passwords and session tokens. AI agents don't work that way. They authenticate with tokens, API keys, and certificates that frequently never expire and that carry access indefinitely. Treating an agent like a service account with a standing credential is the wrong model for a system that's making autonomous decisions dozens or hundreds of times a day.

CoSAI's Agentic IAM paper formalizes the alternative: agents as first-class identities, distinct from both human users and traditional service accounts, each one bound to a unique identity tied to verifiable claims about its code, its model, and the environment it runs in.

The permission rule that follows is straightforward to state and easy to violate in practice. When Agent A calls Agent B, Agent B should not automatically inherit everything Agent A was allowed to do. Permission has to be explicitly granted, and it has to narrow at every hop rather than ever expanding past what the agent making the call was authorized for. This is the direct countermeasure to the confused deputy failure described earlier: the Asana incident happened because an intermediary used credentials it had no business using on another tenant's behalf. A system where permission scope can only shrink as it passes between agents closes that specific failure mode.

To make this work in practice, every agent-to-agent call carries a signed delegation token, scoped to the exact task at hand, so the downstream agent never receives more authority than the task requires. That keeps privilege from escalating through the chain and leaves a record that can actually be audited after the fact.

Avoid passing an upstream caller's OAuth token straight through to the next service in the chain. CoSAI recommends performing a token exchange at every trust boundary instead, using RFC 8693, so each hop gets a new token scoped to its specific operation, with its own auditable exchange record, and the original user's token never reaches a downstream service that doesn't need to see it.

Credential architecture that ensures agents never hold the keys they use

Identity and permission scoping only hold up if the credentials behind them are built the same way. The core rule here is that agents should never hold the sensitive credentials they use to act. A stolen token in a conventional application has a contained blast radius, usually one session, one user, one system. In an agentic system, that same stolen token can traverse multiple servers and authorize an entire chain of downstream operations; credential brokering has to happen at the infrastructure level rather than inside each agent.

The vault-based isolation model handles this by keeping every external API credential stored server-side in an encrypted vault. Agents never receive the raw API key at all. Instead, they get short-lived tokens scoped tightly to the task in front of them. Peta, which markets itself as "1Password for AI agents," builds this pattern by pairing an MCP gateway with an agent-facing vault and a policy engine, which is a useful concrete illustration of the architecture rather than the only way to build it.

Just-in-time credential provisioning extends the same logic to timing. Credentials get minted the moment you need them and expire the moment the task finishes. That removes standing access entirely; standing access is what makes a stolen credential immediately useful to an attacker. A credential that already expired by the time someone tries to misuse it is a credential that failed to help them.

A policy engine has to sit in front of credential issuance itself. High-risk actions need a check at the moment credentials get released, not only when the agent first authenticates, because authentication happening once at the start of a session says nothing about whether the specific action an agent wants to take later should be allowed.

Give each MCP server connection its own distinct credential instead of a shared token, so a compromise stays isolated to a single server instead of spreading across the agent's full tool set. And the vault itself needs to function as the authoritative registry of every machine identity in the organization. Every new AI tool a team connects through MCP adds another machine identity that conventional IAM tools were never built to track, and that registry is the only place equipped to see all of them at once.

The MCP gateway as a mandatory control plane between agents and tools

None of the identity and credential principles above hold together without a central point where they actually get enforced. Without a dedicated MCP gateway acting as that control plane, each agent ends up managing its own configuration and its own credentials independently, and no single system can say which tool ran on whose behalf at any given moment.

The gateway sits between the agents and the MCP servers they call, and it's the place where authentication, tool-level authorization, credential brokering, and audit logging all get enforced on every single tool call, not just at the start of a session. That turns organizational policy from something written in a document into something the system mechanically enforces.

Authorization needs to happen at the level of the individual tool call. A team cleared to use a CRM's MCP server shouldn't automatically get access to every tool that server happens to expose. Tool-level access control lists are what make that distinction possible, letting an organization grant narrow, specific permissions instead of an all-or-nothing connection.

The category is already taking shape in the market. Kong launched an MCP Registry inside Konnect in early 2026, built to catalog and govern approved MCP servers. That kind of registry, paired with a gateway enforcing policy on every call, is what makes it possible for an enterprise running thousands of agents to actually answer the question this whole piece keeps coming back to: who did what, through which path, and under whose authority.

Sources

  1. AI Agent Security and MCP Defense Guide
  2. Model Context Protocol (MCP): Security Design ...
  3. Threat Advisory: MCP Threats
  4. After RSAC™ 2026: The MCP Security Question Everyone Kept Asking - Coalition for Secure AI
  5. Model Context Protocol (MCP) Security

More in Agentic AI Foundations