LetterMCP

EU AI Act Compliance Checklist for Agentic Systems

Agentic systems trigger EU AI Act obligations companies aren't yet tracking.

Correspondent · · 11 min read
Cover illustration for “EU AI Act Compliance Checklist for Agentic Systems”
Agentic AI Standards & Frameworks · October 4, 2026 · 11 min read · 2,369 words

A compliance team that trusts its model registry as proof of EU AI Act readiness is already behind. Most programs track which models are in use and under what license. They don't track what autonomous agents built on those models actually do once deployed: which tools they call, which systems they touch, which people their actions affect. That gap is structural. The AI Act's risk framework was finalized before agents with persistent memory, tool use, and multi-step planning became a normal part of enterprise software, so the law was never written with a line item for "agent" in mind.

The Act defines "AI system" in Article 3(1) in functional, technology-neutral terms, on purpose, so that the Regulation wouldn't need constant rewriting every time the underlying tech changed. Nannini et al. (2026) establish that no normative category of "agent" appears anywhere in the Regulation, nor in the draft harmonised standards under Standardisation Request M/613. That omission means the law reaches agents through the same functional test it applies to everything else, while the guidance, tooling, and audit practices built around the Act still assume a model-centric world.

As of early 2026, the AI Office has published only preliminary material addressing agents, autonomous tool use, or runtime behavioral change. The AI Act Service Desk FAQ states that its thinking on agents remains "only preliminary. Nannini et al. Nannini et al. go further: high-risk agentic systems whose behavior drifts in untraceable ways cannot currently satisfy the Act's essential requirements. The starting task for any program serious about this is building an exhaustive inventory of what each agent actually does externally, rather than sorting architectures into categories: the actions it takes, the data it touches, the systems it connects to, and the people its decisions reach.

Agentic systems create compliance obligations the Act already covers but programs aren't tracking

The Act doesn't need an "agent" category to regulate agents. Its existing articles on risk classification, transparency, logging, and human oversight apply by virtue of what a system does, not what it's called, and those obligations stack further once agents start operating in chains or taking consequential actions in the real world.

Article 9's risk management requirements apply to autonomous agents handling financial transactions, medical decisions, or legal submissions, which are likely to count as high-risk under the Act. That triggers conformity assessment, human oversight, and auditability duties. In a chain of agents, Nannini et al. are specific that the compliance boundary extends to every agent performing a high-risk function in that chain, not just the one that started it. An orchestrator that looks low-risk on its own can still sit upstream of three or four agents each carrying full high-risk obligations.

Article 50's transparency duties take effect from 2 August 2026, and from that date, providers must design any system meant to interact directly with people so users know they're talking to AI. Sidley's analysis confirms these duties mostly survived the Digital Omnibus delays intact. The one carve-out: the content-marking obligation under Article 50(2), for systems already on the market before 2 August 2026, got pushed to 2 December 2026 under the Digital Omnibus on AI, agreed provisionally on 7 May 2026 and in force since 27 July 2026. The Commission's draft Guidelines confirm agentic systems fall inside Article 50 whenever their output is meant to be seen directly by a person. An agent that only talks to another machine sits outside that specific disclosure rule, though its high-risk obligations, data protection duties, and content-marking rules can still apply regardless.

Article 12 requires high-risk systems to support automatic recording of events across their operating life, so emerging risks can be caught and post-market monitoring can actually function. Deployers have to keep those logs for at least six months under Article 26. Article 14 requires that high-risk systems let a human actually oversee and intervene in what the system does while it runs. The oversight has to be built into the agent's workflow from the start, not bolted on after deployment.

None of this runs in isolation. An agent provider is also managing duties under GDPR, the Cyber Resilience Act, NIS2, the Digital Services Act, the revised Product Liability Directive, and whatever sector rules apply on top. Nannini et al. offer the first systematic map that lays all of these out together alongside the AI Act, rather than treating each regime as its own silo. And compliance here isn't a document that gets filed once a year. A KPMG Q4 2025 AI Pulse Survey found 75% of large-enterprise leaders name security, compliance, and auditability as the most critical requirements for agent deployment. A checklist with every box ticked is not the same thing as a control a regulator or an insurer can test on any given Tuesday.

The hardest piece of this to govern is behavior that changes after deployment. Kaptein et al. (2026) show that agent behavior is non-deterministic and depends on the path it takes through a task, in ways no design-time review can fully anticipate. System prompts and static access rules only shape or restrict a subset of the paths an agent might take. Evaluating the agent's actual behavior at runtime, path by path, is the only approach general enough to cover every obligation that depends on what the agent actually did, not just what it was told to do.

MCP as the largest single compliance exposure in most enterprise agent deployments

The Model Context Protocol has become the wiring that connects agents to tools, data, and each other across most enterprise deployments.

Consider the shape of a typical tool-poisoning incident. A tool gets approved through normal review. An agent's data query inherits the permissions of the analyst who set it up. The outbound call goes out to a server already on the allowlist. Every individual step passes inspection, and the system still gets compromised, because the vulnerability sits in the trust boundary itself. MCP mixes tool descriptions (instructions) in with the data an agent processes, so a quiet edit to a tool's metadata can redirect what the agent does just as effectively as rewriting its system prompt. No model-level audit catches that, because the model never changed.

Every agent added to an enterprise stack also creates a new non-human identity that needs API access and machine-to-machine authentication, a problem legacy identity management systems weren't built to handle. That compounds the Article 12 logging duty with a second, harder problem: tracing which identity did what, across a chain of agents and tools, after the fact. Nannini et al. frame AI agent security in 2026 as a supply chain problem first and a prompt injection problem second, with MCP running as the connective tissue across nearly every major incident category.

That framing changes what a checklist needs to cover. It's not enough to inventory models or even agents in isolation. The inventory has to capture every tool connection, every inherited permission, and every machine identity an agent creates, because that's where the actual exposure lives.

Step 1, Build the agent inventory before classifying anything else

No article of the Act can be applied to an agent that hasn't been mapped first, so the inventory comes before classification, before risk assessment, before anything else on the checklist. It is a separate exercise, not a model registry with a few agent names tacked onto the bottom, that captures what each agent actually does once it's running.

For every agent in the estate, the record should capture its intended job, who uses it, and who could be affected by what it does. It should capture what data the agent can reach, what actions it's able to take on its own, what tools it calls, and what systems those tools connect it to. It should also capture where the agent's outputs end up, including whether those outputs reach anyone in the EU, since that question determines scope even when the organization deploying the agent is based somewhere else.

An agent's classification can shift the moment its job changes, even without a single line of its code changing. A document summarizer connected to hiring records and asked to rank job applicants is no longer doing the same regulatory job it started with, even though nobody touched the underlying model. That means the inventory has to track what the agent is actually doing in production.

In a chain of agents, the inventory needs to cover every agent performing a function with real consequences. The orchestrating agent is often the one a team remembers to document. The three or four agents it calls downstream, each taking its own consequential action, are just as much inside the Act's reach.

Each agent's entry should also record which policy version was in force at the time it was deployed. Runtime behavioral drift, the point where an agent's adaptive behavior crosses from expected learning into what Article 3(23) treats as a substantial modification, is itself a compliance event. Catching that drift requires a versioned baseline to measure against, so the inventory has to be a living record updated as agents change. Kaptein et al. formalize this same idea from the runtime side: governing an agent means tracking its identity, the partial path it has taken through a task, the next action it proposes, and the state of the organization around it. The inventory is the starting data for exactly that kind of ongoing evaluation.

Step 2, Classify each agent's risk level against Annex III before applying any other obligation

Classification depends on what an agent does and who it affects, not on how capable the model underneath it happens to be. That distinction trips up more compliance programs than almost anything else in this process, and it needs revisiting every time an agent's purpose, data access, tools, or user base changes.

Running a powerful model does not automatically make an agent high-risk. An agent that summarizes internal documents for an internal team may carry few obligations under the Act at all, even if it's built on the most capable model available. An agent that screens job applicants or decides who gets access to an essential service is in Annex III high-risk territory and needs a full conformity assessment, regardless of which model powers it. The risk lives in the decision the agent makes.

Autonomous agents that take consequential actions directly, moving money, making medical calls, filing legal submissions, are likely headed for high-risk classification. That brings Article 9's risk management duties, Article 12's logging requirements, Article 14's human oversight rules, and a conformity assessment, all at once. Teams sometimes assume that giving an agent more autonomy reduces risk, on the theory that fewer humans in the loop means fewer human errors. The Act's structure runs the other way: more autonomy over a consequential action means more obligation, not less.

Two misclassification patterns appear constantly in how teams apply this framework. One is assuming that because the underlying model has already gone through some compliance review, the agent built on top of it inherits that clean bill of health, and it doesn't. The model's certification says nothing about what a specific agent configured on top of it actually does in production. The other is assuming a system is low-risk by default unless proven otherwise, when the Act's structure places the burden the other way: the agent's actual function has to be checked against Annex III before anyone can conclude it falls outside that category.

Before spending effort on any of this, check whether the agent's use case is prohibited: certain kinds of harmful manipulation, social scoring, or emotion recognition in workplaces and schools sit in that bucket. A prohibited agent doesn't need governance. It needs to be shut down.

In a multi-agent chain, classification has to happen agent by agent. Each one performing a high-risk function carries the obligations that attach to that function, no matter how the broader system architecture gets described on a slide deck. Nannini et al. lay out a taxonomy of nine agent deployment categories that map concrete actions onto the regulatory triggers they carry, offering the most detailed public mapping of this kind available. And none of this depends on geography in the way teams sometimes assume: a provider or deployer based outside Europe is still inside the Act's scope the moment an agent's output reaches someone in the EU.

Step 3, Apply Article 50 transparency obligations, with exact deadlines, to every agent that interacts with users

Diagram: Article 50 Transparency Deadlines: What Changed and What Didn't. Visualizes: Show the two-date compliance timeline for Article 50 transparency obligations, contrasting what was delayed versus what remained intact.

Article 50's transparency duties took effect on 2 August 2026 for most interactive AI systems, and the Digital Omnibus delay touches only the content-marking obligation for systems already on the market before that date. Any organization that hasn't already built disclosure into new agent deployments is out of compliance right now, not at some point in the future.

The core provider duty is straightforward to state and easy to underbuild: any AI system meant to interact directly with people has to be designed so those people know they're dealing with AI, unless that fact is already obvious from context. Sidley's analysis is specific about timing here: the notice has to appear no later than the first interaction, not somewhere buried in a settings page a user might never open.

The Commission's Guidelines confirm that agentic systems fall inside Article 50 wherever their output is meant to be seen directly by a person. The Commission names chatbots as an example and confirms agents generally fall within scope. AI avatars, meaning digital people that speak and respond to a user, are covered the same way.

Agent-to-agent interaction works differently. An agent that only talks to another machine, with no output a person directly perceives, sits outside the specific notice requirement that Article 50 attaches to direct human interaction. That doesn't clear the agent of every other duty. High-risk obligations under Article 9, data protection requirements under GDPR, and content-marking rules can still apply to that same agent, depending on what it's doing and what data it's touching along the way. Mapping which agents in the estate interact with people, versus which ones only talk to other systems, is the final piece of the checklist, and it determines exactly which notice obligations apply, and when, to each one.

Sources

  1. AI Agents Under EU Law A Compliance Architecture for AI Providers
  2. EU AI Act Transparency Obligations: Preparing for Compliance by 2 August 2026
  3. Runtime Governance for AI Agents: Policies on Paths
  4. Frequently Asked Questions
  5. Transparency obligations under Article 50 of the AI Act

More in Agentic AI Standards & Frameworks