MCP Tool Description Best Practices for Enterprise Servers
Tool descriptions shape model behavior and need security review like code does.

A tool description in the Model Context Protocol is not a comment left for a human to skim; it is the primary channel through which a language model learns what a tool does, when to call it, and what to hand it, and it functions as an instruction rather than a label. The architecture makes this unavoidable: a host passes the model a list of available tools along with their descriptions, and the model reads that list and decides, on its own, which tool to call and what arguments to send it. The model drives execution. The user doesn't choose the tool call; the description does, by shaping what the model believes the tool is for.
That's different from a docstring sitting in a codebase, read once by an engineer and then forgotten. A tool description gets re-read by the model every time it's relevant, and it shapes the reasoning that precedes every single tool call made against it. Nothing about this stays confined to the prose in the description field, either. Parameter names, enums, default values, the whole JSON Schema attached to a tool, all of it is readable metadata in the model's context. The full schema, not just the sentence a developer wrote to describe the tool, is the instruction surface.
This produces a trust problem with no easy patch. The Cloud Security Alliance's Agentic MCP Security Best Practices guide names the gap between the LLM and the MCP client as the first trust boundary in the whole system. The model reads a tool description and builds an invocation from it, but it has no independent way to confirm that description is accurate, or that nobody has altered it since the last time the model saw it. Once that's clear, the security implication follows on its own, without needing dramatic language to make the point: anything an author can put into a description can shape a tool call, so writing one is a security-relevant act, whether the author treats it that way or not.
The protocol's rapid growth and the governance vacuum around descriptions
MCP spread through the industry faster than most protocols do, and the tool description surface grew right along with it, largely unexamined. OpenAI adopted MCP in March 2025, Microsoft announced support for it in Copilot Studio that same month, and by late 2025 combined SDK downloads had climbed into the tens of millions per month. That is adoption at a pace most security practices never get to match.
The incidents from 2025 show what happens when practice lags behind adoption. Cross-tenant data exposure hit Asana. Prompt injection attacks were demonstrated against the GitHub MCP server. Anthropic's own MCP Inspector developer tool carried an unauthenticated remote code execution flaw. Malicious npm packages compromised the supply chain feeding into MCP deployments. None of these trace back to a single bad actor finding one clever trick. They trace back to a protocol that got adopted by major platforms within the same month, while the discipline around writing and reviewing the artifacts that drive model behavior stayed informal.
The scale of that gap is measurable. Research published in 2026 found that only a small fraction of tool descriptions in active use satisfy all five recognized quality components at once. The overwhelming majority of deployed descriptions are incomplete, ambiguous, or structurally unsound in some way. The July 2026 spec revision hardened authorization substantially, closing real gaps at the protocol level. The gap that remains sits downstream of the spec, in how teams actually write descriptions day to day. Most organizations still treat a tool description as a writing task, something a developer drafts quickly and nobody reviews with the rigor applied to the code behind it. Spec hardening doesn't fix that. Only a change in practice does.
The three ways a tool description becomes a weapon
Three distinct mechanisms let an attacker turn a tool description against the system it's supposed to serve, and all three share one trait: a human operator skimming a tool list in a client UI won't see any of them, while the model reading the full description sees everything.
The first is tool poisoning and schema poisoning. A 2025 analysis of open-source MCP servers found that 5.5% showed signs of tool poisoning, and that carries the risk of over-privileged access. A description might legitimately claim to search local files, while a hidden clause buried in that same description instructs the model to also read a file like ~/.aws/credentials and smuggle its contents out through a parameter that looks unrelated to credentials at all. Schema poisoning extends the same idea past the prose description, into parameter names, enums, and default values: the entire JSON Schema counts as attack surface, not just the sentence a reviewer might glance at. Most client interfaces only show a user the tool's name. The model sees the complete description. A 2026 peer-reviewed security analysis traced this back to a protocol-level gap: there's no mechanism for capability attestation, so a server can claim whatever permissions it wants and the model has no way to check the claim against reality.
The second mechanism is the rug pull. A server behaves exactly as expected during the review that got it approved, then a later update, whether pushed automatically or silently re-fetched at runtime, swaps in different, malicious behavior. Because most clients don't diff tool definitions between sessions, a server can change what it does without anyone re-reviewing it. A 2025 demonstration against a WhatsApp MCP server showed how this plays out in practice: a second, malicious server connected to the same agent used poisoned tool descriptions to direct the agent to pull message history out through the trusted WhatsApp server and send it to an attacker.
The third is cross-server shadowing combined with indirect prompt injection through tool results. Here, a malicious server's description doesn't attack its own tool calls directly. Instead it instructs the model to change how it uses a different, trusted server's tools, targeting the model's behavior toward a server it has every reason to trust. A 2025 security analysis on arxiv points to bidirectional sampling without origin authentication as the protocol-level weakness that lets this kind of server-side prompt injection happen.
What the MCP specification requires tool descriptions to contain
The specification sets a structural floor for tool descriptions, and most servers published today fall short of clearing even that floor; spec compliance alone doesn't solve the problem. The current spec, dated 2026-07-28, defines a tool object with a required name, an optional human-readable title, a description field, a required inputSchema that must be valid JSON Schema, and an optional outputSchema. Tool names are capped at a maximum character length and restricted to ASCII letters, digits, underscores, hyphens, and dots, with no spaces or special characters allowed.
That covers the shape of a tool object. It says nothing about whether the description inside it is actually good. A 2026 rubric published on arxiv lays out five components a quality description needs: Purpose, Guidelines, Limitations, Parameter Explanation, and Length or Completeness, with a recommendation of at least three to four sentences to give the model enough to work with, more for complex tools. Against that rubric, only 2.9% of tool descriptions in the study satisfied all five components together. Nearly every description examined was missing something the model needed.
The spec also draws a line between two different places guidance can live. Server-level instruction blocks are where workflow guidance, constraints, and usage conventions belong. Tool-level descriptions are meant to describe what a specific tool does, not to carry broad instructions about how the agent should reason about using it. Collapsing that distinction, stuffing workflow guidance into individual tool descriptions, is one of the quieter ways servers fall short of the floor the spec actually draws.
Writing tool descriptions that are narrow, honest, and structurally safe
Good tool description practice isn't a matter of clean writing for its own sake. Every choice that makes a description less ambiguous also shrinks the room a poisoning attack has to work in, and that connection runs through every recommendation below.
Start with naming. A tool called list_open_issues(repo) or create_comment(issue_id, body) tells the model, and anyone reviewing the server, what can happen when it's called. A tool called run_github_api(method, path, body) tells nobody anything, because it can do almost anything. The same logic applies to get_customer(id) against run_sql(query). Generic shell, SQL, or HTTP tools turn a single successful injection into full compromise, since the tool itself places no limit on what the injected instruction can accomplish. A narrower tool keeps the blast radius of any single poisoned or injected call small by construction, not by hope. A tool built to do everything gives the model no way to tell its legitimate uses from its dangerous ones, since nothing in its shape separates them.
Keep the language in a description honest and purely functional. Because the description sits inside the model's context, it's part of the attack surface in exactly the same way the parameter schema is, so the safest descriptions stay narrow, factual, and free of anything that reads as an instruction to the model beyond a plain account of what the tool does. A description that tells the model how to behave, rather than simply what the tool is, opens the same channel an attacker would use to inject instructions of their own; from the model's point of view, there's no way to tell the two apart. Workflow guidance, constraints, and usage conventions belong in the server-level instruction block, not the individual tool description.
Make parameters explicit, with real types attached. A description that mentions a parameter only in passing, in a sentence of prose, without defining it as a proper argument with a data type, leaves the model guessing when it tries to build a call, and guessing produces broken calls. Replacing that kind of vague mention with an explicit argument, and specifying the expected format directly (a date field marked yyyy-mm-dd, for instance), fixes a usability problem and a security problem at the same time, since an implicit, undefined parameter is exactly the kind of gap schema poisoning exploits. A 2026 enterprise checklist from Maxim AI treats every argument the model composes as untrusted input by default, a posture that only works if the schema is tight enough to catch a malformed or injected argument when it appears.
Watch the parameter count. AWS Prescriptive Guidance for MCP Strategies recommends keeping each tool to around eight parameters or fewer. Past that point, the model's ability to compose a correct call starts to degrade, and at the same time the schema itself grows large enough to give a poisoning attempt more places to hide.
Finally, account for state directly in the description. Where a tool creates a handle that outlives the single connection that made it, the retention policy for that handle belongs in the creation tool's own description, something as simple as stating that a basket expires after 24 hours of inactivity. Leaving that detail out doesn't make the state disappear. The model creates persistent state without knowing it has, which produces unpredictable behavior later and leaves an audit trail with a gap in it.
Agent failures, governance failures, and security failures share the same root cause
The same defects in a tool description that open a security hole also break the agent's own reliability and make governance audits unenforceable.
On the reliability side, a poorly written description degrades the model's reasoning directly: the model makes worse choices, calls the wrong tool, or fills in the right tool with the wrong parameters, and each retry that follows adds more to the context window, compounding the original problem instead of resolving it. Tools that resemble each other too closely, a tool list with too many options, or names that don't clearly separate one tool's purpose from another's, all push the model toward exactly this kind of confusion, because nothing in the descriptions gives it a sharp line to draw between them. What follows is a loop: the agent retries, the context grows, reasoning gets worse with each pass, and confidence in the whole system erodes, often well before anyone traces the failure back to a vague description sitting at the root of it.
On the audit side, the gap in governance is visible as missing information in the log, not as wrong behavior. A traditional audit log records what data got touched. Governing an agent requires something more: a record of why the agent touched it and what it decided to do as a result. A complete log entry needs the agent's unique identifier and version, the specific permissions delegated for that execution, the tool actually invoked, the governance policy decision that was made, and the reasoning step the agent generated right before acting. When the tool description behind that action was ambiguous, or carried instructional language instead of a plain account of function, the reasoning the agent logs inherits that same ambiguity. The audit trail ends up full of noise where intent should be, and no amount of log retention fixes a record that was never precise to begin with.
A model reads an informative sentence and a manipulative sentence the same way, with no way to tell them apart. The description is the interface, the interface is executed, and writing one carelessly is a decision with consequences in security, in reliability, and in the audit record, all at once.
Sources
- Agentic MCP Security Best Practices Guide
- MCP Security: Top 7 Risks and Critical Best PracticesMCP Security Best Practices: The Complete 2026 Guide
- MCP Security Best Practices: Enterprise Checklist 2026
- MCP Security Best Practices: A Practical Guide for 2026
- Understanding Model Context Protocol Security (MCP) in 2026
- Model Context Protocol strategies on AWS - AWS Prescriptive Guidance


