Vendor Detection Coverage for MCP Threat Categories
Vendors catch some MCP threats but miss critical gaps in supply chain and tool poisoning attacks.

The Model Context Protocol was built so AI agents can talk to tools and data sources, without custom integration work for every pair [1][2][3]. Security was never the design goal, and that gap now sits squarely on the desk of every enterprise security team running agents in production. This piece maps which MCP threat categories current vendor tools actually catch, which they catch halfway, and which slip through entirely, because that map is the starting point for building a posture that holds up.
MCP's design and the threat surface existing security categories don't map onto
MCP leaves authentication optional. It defines a narrow set of access control requirements, OAuth 2.1 for remote servers and role-based authorization added in 2026 updates, but authorization itself is optional, so enforcement is left to whoever builds the implementation. Logging gets similarly thin treatment: the spec offers basic guidance and nothing more, so each team decides on its own whether to build real audit trails or skip the work.
The protocol runs on two transport mechanisms, STDIO for local integrations and Streamable HTTP for remote ones, and each one opens a different path for attack. Neither one maps well onto the perimeter models or API-gateway models that security teams already have running. A firewall rule or a gateway policy built for REST traffic doesn't know what to do with a local process spawned over STDIO, and it often can't see inside an HTTP stream carrying MCP's own message format.
Cloud Security Alliance's research note lays out how bad this gets in practice: MCP components, meaning clients, servers, proxies, and hosts, can be configured to access or process data with no required access control measures, and plenty of real deployments skip authentication. Add to that the way MCP servers work in most enterprise setups: one server often holds credentials for several different systems behind a single interface. A compromised server is a single point of failure across everything that server touches, which is a risk profile far more dangerous than a typical API breach.
One more structural gap makes this worse. The protocol has no built-in way to track changes to a tool's definition, and no requirement to re-approve a tool after it changes. A tool an agent trusted yesterday can look identical today and behave completely differently, with no signal reaching either the agent or the security team watching it.
The formal threat taxonomy security teams are now expected to cover
OWASP's MCP Top 10 project gives the field a shared vocabulary for what can go wrong: token mismanagement and secret exposure, privilege escalation through scope creep, tool poisoning, software supply chain attacks, command injection, intent flow subversion, insufficient authentication and authorization, lack of audit and telemetry, shadow MCP servers, and context injection. Ten categories, each demanding a different kind of detection, and that breadth is itself the problem. A team with strong coverage against supply chain attacks can have close to zero visibility into tool poisoning, because the two require completely different tools to catch.
Tool poisoning deserves a closer look, because it exploits something basic about how agents work. An agent reads a tool's description to decide how and when to call it, and it extends the same trust to that description that it gives its own system prompt. A malicious server can hide commands inside that description text. The model follows them: it reads files it shouldn't, leaks secrets, and still hands back an answer that looks completely normal.
Microsoft's Defender research team disclosed a case in mid-2026 that shows how this plays out. A finance team's agent connected to a third-party invoice enrichment tool, one that had been approved but never actually given a real security review. The attacker later updated the tool: the visible name and summary stayed the same, but a hidden instruction was buried inside the description text. Because MCP picks up description changes automatically, the poisoned version went live immediately, with no re-approval step to catch it.
Rug-pull attacks work the same way but stretch over a longer timeline. A malicious server can present harmless tools at first, earn approval, and then quietly change its tool definitions or behavior later. The agent keeps transacting with that server session after session, and nothing in the protocol flags the change.
The Agentjacking incident involving Sentry's MCP integration, disclosed in June 2026, shows what happens when even the vendor response falls short. Tenet Security found a large number of organizations running injectable DSNs, and in testing, the attack succeeded against agents most of the time. Sentry declined to fix the root cause and instead shipped a filter that targeted one specific payload string. The detection gap on the client side and the gateway side stayed wide open.
Supply-chain and credential-sprawl threats, which enter before runtime detection can see them
OX Security's research found a command execution flaw built directly into Anthropic's official MCP SDKs, across Python, TypeScript, Java, and Rust. The STDIO transport takes incoming configuration and passes parameters straight to the host operating system's shell, with no input sanitization. Any process command can execute this way, whether or not it ever sets up a valid MCP server. The flaw touched an enormous number of vulnerable instances, because it spread across a supply chain built on a huge volume of package downloads. Anthropic confirmed the behavior was intentional and chose not to change the protocol's architecture, which leaves the fix in the hands of every downstream developer using the SDK.
Credential sprawl is the default architecture, not a side effect of bad configuration. Setting up an MCP server typically means writing credentials, API keys, OAuth tokens, database passwords, straight into a plaintext config file. That file sits readable by the same AI agent the server just handed a set of tools to.
The Wiz Research disclosure involving Amazon Q's VS Code extension shows how this plays out in a real tool. A high-severity flaw, CVE-2026-12957, let a malicious repository achieve arbitrary code execution and steal cloud credentials through a crafted workspace MCP configuration file. A second flaw, CVE-2026-12958, patched a separate symlink-validation issue. Because the MCP server processes that got spawned inherited the developer's full environment, exploiting the flaw handed an attacker immediate access to cloud infrastructure, not just the local machine.
Where the three dominant vendor defense approaches stop working
Vendors selling MCP security tools today cluster into three camps: allowlisting, gateway routing, and runtime inspection. Each one covers a real part of the threat surface, and each one has a blind spot the other two don't fill.
Allowlisting controls which MCP servers an agent is permitted to connect to. It doesn't look at what's inside the traffic. A poisoned tool description coming from a server that's already on the allowlist sails right through, because the control only checks identity, never content. Allowlisting also fails to catch shadow MCP servers unless every single agent in the environment is forced to go through the same control plane, with no exceptions.
Gateways handle authentication and routing well, for any agent that actually uses the gateway. An agent that calls an MCP server directly, bypassing the gateway, is invisible to it. Most gateways also stop short of deep content inspection on responses coming back from a server. A gateway enforces who is allowed to talk to what and where that traffic goes. It says nothing about what's actually inside the message.
Runtime inspection catches the threats the other two miss: rug-pulls, response injection, credential leaks happening live in traffic. Pre-deploy scanners catch static tool poisoning before anything ever runs. But a pre-deploy scanner has no visibility into runtime behavior, and runtime inspection doesn't replace authentication or network allowlisting. Each tool does its job and stops exactly where the next one needs to begin.
Some security teams argue a gateway-first posture is enough on its own, since it centralizes authentication and traffic routing in one place. The NSA's advisory on MCP documents authentication failures, trust boundary vulnerabilities, and insufficient controls as ongoing risks across agentic AI deployments, and treats the agentic environment as something that has to be defended as a continuum. Organizations running agents against sensitive enterprise data need both layers: a gateway to control who connects and where, and runtime inspection to see what's actually moving through that connection.
The authorization and audit-log gaps that make incident response unreliable even when detection fires
Detection generating an alert doesn't mean a security team can act on it. Authorization and audit logging are the two gaps behind that problem, and together they mean a signal can fire with nobody able to say what the agent was allowed to do or what it actually did.
The updated MCP specification from July 2026 tells a server which client is calling it and which issuer vouched for that client, and goes no further than that. It doesn't define fine-grained, per-tool action policies beyond basic OAuth scopes and audience binding, so whoever implements the server has to make every harder call on authorization. Authentication answers who is calling. Authorization answers what they're allowed to do once they're in, and that second question is where most agent incidents actually begin.
Audit logging compounds the same problem from a different angle. Every gateway and every inspection tool on the market generates logs, but none of them agree on what fields to include. OWASP names lack of audit and telemetry directly as MCP08 on its top-10 list, and there's still no shared schema that incident responders across different tools can work from. A security team pulling logs from three different products during an active incident is reconciling three different formats before it can even start reconstructing what happened.
The specification's own guidance on logging stays basic, so each implementer has to decide what comprehensive auditing looks like. In practice that means many deployments either skip logging altogether or record only minimal operational metadata, just enough to confirm a call happened, not enough to show what the call actually did once the agent was inside a connected system. An agent that moves laterally from one connected tool to another during an incident can leave a trail too thin to reconstruct, turning a detected breach into an open question about scope.
What a threat-category-aware MCP security posture requires
Detection capability varies sharply from one threat category to the next, so a security posture built around a single vendor's tooling will always leave gaps somewhere in the taxonomy. The right approach starts with the threat categories themselves and treats tooling as the second decision, not the first. Teams that map their coverage gaps before buying anything end up with fewer blind spots than teams that buy a gateway and assume the problem is solved.
Supply-chain threats need pre-deploy scanning of MCP server packages and their configuration files, registry vetting before anything gets approved, and credential rotation policies that run on a schedule. These controls operate before an agent ever makes its first tool call, and no amount of runtime monitoring substitutes for work that has to happen earlier in the pipeline.
Shadow MCP servers and unauthorized connections need network-level visibility and an inventory control plane that every agent in the environment is required to pass through, with no side doors. A gateway that an agent can simply route around moves this gap somewhere the security team can't see.
Authorization gaps need policy enforcement down at the level of individual tool calls, not just a checkpoint at the perimeter. Knowing which server an agent connected to tells a security team far less than knowing what that server was permitted to do once the connection was live.
Audit and incident response need tamper-proof, structured logs, and every single action an agent takes needs identity attached to it. Aggregating logs from several disconnected tools, each with its own schema, won't get you there. It delivers a reconciliation problem on top of the original incident.
Lyzr's September 2026 guide states the architectural implication directly: registries need to connect into a broader AI control plane that governs agents, identities, and runtime activity together. A stack built from a standalone agent builder, a separate gateway, and a bolted-on security scanner isn't a strategy, it's three tools that each cover their own slice of the taxonomy and leave the seams between them exposed.


