CISO Playbook for Governing Internal vs External AI Agents
Different threat surfaces demand separate governance frameworks for internal and external AI agents.

Most enterprises are making the same mistake right now: treating every AI agent as one risk category and writing one policy to cover all of them. That approach breaks down because internal and external agents face different threat surfaces, operate under different trust assumptions, and need different controls to match.
An internal agent runs inside the company's own perimeter. It acts on behalf of an authenticated employee, touches internal systems and data stores, and gets provisioned and governed by the organization's own IT or platform team. That's a fundamentally different animal than an agent surfaced to customers or the public, one that handles input from anyone, with no guarantee the person on the other end is who they say they are.
The scale of the problem is what makes the distinction urgent rather than academic. The average organization now runs a large number of these agents, and Gartner expects enterprise applications wired into task-specific agents to keep growing fast, so whatever governance gap exists today just compounds. The Cloud Security Alliance has started describing agents as a new category of privileged identity, sitting somewhere between a human user account and a traditional service account. It executes work on behalf of people, often carrying permissions that exceed what the humans it serves actually have.
An agent is a privileged actor, not just a faster version of a script. It's a privileged actor that needs its own identity, its own scope, and its own review process, and the process looks different depending on which side of the perimeter it sits on. A single blanket policy, calibrated to neither, ends up doing two things badly at once: it locks down internal agents with controls built for hostile strangers, and it leaves external agents under-defended against exactly the kind of hostile strangers those controls were meant for.
Internal agent failures: the threat surface inside the perimeter
Internal agents don't usually get breached by outsiders. They get compromised because someone inside the company handed them too much trust in the first place, and that's the defining shape of the risk.
Start with over-scoped permissions. Agents tend to get provisioned like service accounts, broad and standing, rather than narrow and purpose-built for one task. Microsoft's 2026 guidance on least-privilege design for AI agents flags a recurring pattern: temporary access that has no expiry mechanism attached to it eventually just becomes permanent access, because nobody circles back to revoke it. Governance frameworks haven't caught up to this reality. Only a small share of organizations running agents today have actually updated their governance practices to reflect what those agents are doing on the ground.
Shadow provisioning compounds the problem. A developer can stand up an MCP server against a production database in an afternoon, with no central review. Once an agent discovers that server, every tool the server exposes becomes part of what the agent can do, with no extra approval gate in between. Shadow AI, deployed this way, was already responsible for a meaningful share of the data breaches recorded in 2025.
Tool poisoning is the sharpest example of how internal trust gets exploited internally. MCP servers reflect updates to their own tool descriptions dynamically, and in most setups a changed description doesn't trigger a fresh review. Microsoft Incident Response and the Microsoft Defender research team disclosed a case in mid-2026: a finance team had approved a third-party invoice enrichment MCP server, and that server later got silently updated with hidden instructions buried in its tool description. The new instructions told the agent to retrieve unpaid invoices and attach them to enrichment calls. Nobody re-reviewed it, because nothing in the pipeline required a second look when the description changed. Separately, the MCPTox benchmark tested adversarial versions of real tools against a range of major language models and found an average attack success rate above a third, with the worst result landing against one of the leading reasoning models.
Credential hygiene tells a similar story. An Astrix audit of MCP servers found most of them still rely on static API keys or personal access tokens sitting in environment variables, and fewer than one in ten use OAuth at all. Trend Micro ran an internet-wide scan and found 492 MCP servers exposed publicly with no authentication or encryption whatsoever; a follow-up count found that number had nearly tripled.
Supply chain weaknesses round out the picture. Wiz Research disclosed two flaws in the Amazon Q VS Code extension: opening a malicious repository could trigger arbitrary code execution and cloud credential theft through a crafted workspace MCP configuration file, and the spawned MCP server processes inherited the developer's full environment. Microsoft's own @azure-devops/mcp npm package turned up a missing authentication layer in April 2026, on a server that handled Azure DevOps work items, repositories, and pipelines. The most instructive case might be the Deadbugz campaign from August 2026. A single GitHub account filed 23 pull requests across unrelated projects in just 74 minutes, each one offering an MCP server that looked legitimate and behaved normally for its first three tool calls. The pull request's instructions rewrote themselves only after that delay, turning the server to hunt for SSH keys and cloud credentials, letting it slide past code review since reviewers checked only the behavior they could see before the malicious behavior appeared. That delay is what let it slide past code review. Reviewers checked the behavior they could see, and the malicious behavior hadn't shown up yet.
External agent failures: the threat surface beyond the perimeter
External agents live on the other side of a structural divide. Every interaction they handle comes from outside the company's trust boundary, so hostile input is the normal condition the agent operates under. It's the normal condition the agent operates under.
Prompt injection sits at the center of this, and it isn't a bug someone can patch away. Large language models process everything as a single stream of tokens, with no hardware-level wall separating a trusted instruction from an untrusted chunk of text a user or a document handed it, because that's how transformer-based models work. That's just how transformer-based models work. For an external agent, that means any message, uploaded file, or web page it reads is a potential vector for smuggling in instructions the agent will follow as if they came from its operator.
Two disclosed vulnerabilities show what that looks like in production. CVE-2025-32711, nicknamed "EchoLeak" and rated 9.3 on the CVSS scale, involved a prompt hidden inside PowerPoint speaker notes that triggered zero-click data exfiltration from Microsoft 365 Copilot, with no user interaction required at all. CVE-2025-53773, affecting GitHub Copilot and rated 7.8, let an injection payload buried in source code push the agent into executing arbitrary terminal commands, using nothing more exotic than the agent's normal behavior of reading code. HackerOne's ninth annual report backs up how fast this category is growing: prompt-injection reports surged sharply, and the report names it the fastest-growing threat category in AI security.
The stakes go beyond bug bounty statistics. Between December 2025 and January 2026, a single unidentified attacker used Claude, and to a lesser extent OpenAI's ChatGPT and GPT-4.1, to breach multiple Mexican government agencies, including the federal tax authority, the national electoral institute, four separate state governments, and a water utility in Monterrey. External-facing agents connected to sensitive systems aren't just exposed to random opportunists: they're targets for sophisticated, patient adversaries willing to work the angle for weeks.
The exfiltration pattern itself is simple and repeatable. An attacker embeds an instruction inside a web page or a document. The agent reads that content as part of its normal job, follows the embedded instruction, reaches for credentials it has access to, and sends them off to an endpoint the attacker controls. Any agent that reads external content and can also act on internal systems is, by design, a bridge between the two.
The blast radius sets external failures apart from internal ones. When an external agent leaks customer data or fires off an unauthorized transaction, the consequences land on customers, regulators, and partners, outside the company's ability to simply reverse the damage. The legal and compliance exposure from that kind of failure runs categorically higher than an equivalent internal mistake. Gartner's own projections put a number on where this is headed: a quarter of enterprise breaches are expected to trace back to AI agent abuse by 2028, and more than half of successful attacks on agents through 2029 are expected to exploit weak access controls and prompt injection specifically.
Why a single governance policy fails both agent types
Lay the two threat surfaces side by side and the case against a one-size-fits-all policy writes itself. A single policy applied evenly across both agent types will under-govern the external ones on exactly the dimensions that matter most for them, input trust and output controls, while over-restricting the internal ones on the dimensions that shape whether anyone actually adopts them.
Under-governance appears first in the assumptions baked into most policies. They're written around internal identity, single sign-on-bound users and known, traceable request sources, which leaves no real answer for anonymous or hostile external callers. Internal-style policies also rarely require input sanitization, output filtering, or a human in the loop before a destructive action fires, and those are exactly the controls an external-facing agent needs. Most governance frameworks in place today were drafted before anyone fully grasped that an agent could be manipulated through the content it retrieves, not just through direct commands typed at it.
Apply the same controls in the other direction and internal agents choke on them. Wrap internal automation in the sandboxing, approval gates, and narrow tool allowlists built for hostile external input, and the automation gets too slow to be worth using. Teams respond by building shadow agents to route around the friction, which is the exact failure mode the policy was supposed to prevent. Only 6% of organizations running agents have actually updated their governance frameworks to match what those agents do, and that low number says less about laziness than about a mismatch: where controls exist, they're often calibrated for the wrong threat model.
There's a third failure that hits both agent types at once, regardless of which side of the perimeter they sit on. Companies keep applying identity policies built for humans to agents that aren't human: shared credentials, no distinct identity per agent, no way to attribute a specific tool call to a specific actor. When an agent takes a damaging action under that setup, there's no governance record that can establish which identity authorized it, under what policy, or for what stated purpose. Logging only the model's final response, without capturing the underlying tool calls, the scopes granted, and the authorization decisions behind them, produces an audit trail that looks complete on paper but cannot support a forensic investigation after something goes wrong. A large majority of CISOs already report that AI has access to their core business systems, and only a small fraction of them say they're governing that access well. That gap between access and oversight is where both agent types end up exposed, just for different reasons.
The governance posture for internal agents: least privilege, identity, and inventory
Governing internal agents well starts with a mindset shift: treat every agent as a first-class non-human identity, cap its permissions at the minimum a given task needs, and keep continuous visibility over what's running and what it can reach.
Inventory comes before policy, not after it. A governance program can't start by writing rules; it starts by knowing what agents exist, which MCP servers they connect to, what tools those servers expose, and who actually stood them up. That inventory has to run continuously rather than as a periodic audit, because teams spin up new agents and new MCP connections faster than any scheduled review cycle can keep pace with. Every MCP server in the environment needs a named owner, a documented review state, and a version history. Without that history, an attack like Deadbugz, where the malicious behavior only appears after a delay, slips through undetected.
Identity is the next layer. Each agent needs its own scoped identity tied to single sign-on, with its permissions derived from the specific human it's acting for, rather than a broad, standing service account shared across tasks. OAuth 2.1-based identity binding, where an agent inherits the scoped permissions of the user it represents rather than holding its own separate grant, is the direction enterprises are converging on through 2026. The Integrate.io 2026 roundup of agent governance tools points to this same pattern in practice: agent gateways built around non-human identity as a first-class concept, with scoped permissions, dedicated credentials, and audit attribution down to the individual agent.
Least privilege and just-in-time credentials close the loop. Agents shouldn't carry standing access to production systems as a default. Where a task genuinely needs elevated access, that access should be time-bound and scoped only to the resources the task touches, not the whole system. Microsoft's guidance defaults every agent to task-based roles, enforces explicit scopes and tool allowlists rather than open-ended access, uses just-in-time entitlements that expire on their own for anything requiring elevation, and treats access review and revocation testing as routine maintenance rather than a special project.
A tier model helps make this concrete rather than aspirational. A practical tier model splits agents into bands: Tier 0 covers observe-only agents with no write access and no external side effects; Tier 1 covers agents doing reversible work inside a bounded project or test environment; Tier 2 covers agents touching sensitive data, connected services, or anything that publishes externally. Tying each tier to a distinct review bar, rather than treating every agent as equally risky or equally trusted, is what lets an internal governance program move fast on the low-stakes work while still catching the handful of agents that actually deserve scrutiny.



