ISO IEC 5338 Requirements for Enterprise AI Agent Lifecycles
The standard predates autonomous agents and needs translation for tool-using systems.

ISO/IEC 5338 is the most complete lifecycle standard available for enterprise AI systems, covering everything from data acquisition to retirement. But the standard was written before autonomous, tool-using agents existed, and applying it now means translating its requirements clause by clause for a kind of system it never had in mind.
What ISO/IEC 5338 covers
ISO/IEC 5338:2023 lays out a full set of processes for managing AI systems built on machine learning and heuristic methods, from the moment data gets collected through model training, testing, deployment, ongoing operation, and eventually retirement. Every part of an AI system's working life gets its own set of requirements.
The standard sits on top of ISO/IEC/IEEE 15288 and 12207, two long-standing frameworks for systems and software engineering. Building on that base, it lays out both general-purpose and AI-specific processes covering how work gets defined, controlled, managed, carried out, and improved over time. 5338 tries to cover the whole arc.
In practical terms, organizations following 5338 are expected to do several things: mark out clear lifecycle stages and process boundaries for each AI system, apply the right mix of general and AI-specific processes depending on the stage, handle verification and validation and testing with real rigor, keep deployed systems running and watched over time, retire systems responsibly with a documented reason for doing so, and tie all of this back into the broader systems engineering lifecycle laid out in 15288. Each one assumes a team that treats the AI system as something with a beginning, a middle, and a planned end.
5338 also fills in gaps that plain software engineering practice tends to skip over: where did the training data come from, how was the model tested, how is drift tracked once the system is live, and where does a human actually step in to check the system's work. These are good questions, and they were the right questions to ask about the AI systems that existed when the standard was written. 5338's completeness as a process framework is also what makes the job of applying it to a new kind of system so consequential. A standard this complete forces every one of its requirements to be reread, one at a time, against a system that behaves nothing like the one it was built for.
Why agentic AI sits outside the standard's assumptions
5338 carries four assumptions inherited from how AI governance has generally thought about these systems: the edge of the system is the edge of the risk, a human sits somewhere in the review loop, the party affected by the system is whoever receives its output, and the system stays inside the boundary it was assessed against.
Agentic AI doesn't have a clean endpoint the way a model that produces a single output does. An agent reads a tool description, decides to call that tool, acts on what comes back, and may repeat that cycle many times before a person ever sees a result. The "perimeter" 5338 assumes, the edge where the system's behavior can be checked and signed off on, simply isn't where it used to be.
The terminology problem makes this harder to pin down. An enterprise working across borders can run into four different definitions of the same word in four different filings.
None of this means 5338 is broken or should be set aside. Governance standards like 5338 are built to be technology-agnostic on purpose: they focus on management systems, risk processes, lifecycle discipline, and organizational accountability, not on runtime control logic that tells a system what to do in the moment. The requirements are sound. They need a different behavioral model sitting underneath them before they can be applied honestly.
An agent can look completely fine at every single step, each tool call reasonable, each decision defensible on its own, while the full sequence of actions adds up to something nobody would have approved. An agent does.
Translating the standard's data readiness requirements for live tool input
5338's rules around data acquisition, preparation, and provenance were written with training pipelines in mind: where did the data come from, how was it cleaned, can its lineage be traced. For an agent, those same questions don't stop once training ends. That data accumulates after the model has already been deployed, in a part of the system 5338's original data requirements were never built to watch.
The Model Context Protocol, Anthropic's open standard for connecting agents to outside tools and data sources, makes this concrete. An agent reads a tool's description at the moment it decides whether and how to use that tool, and it trusts that description with roughly the same weight it gives its own system prompt. That's a direct line from an unvetted piece of text to the agent's behavior, no training pipeline involved.
Tool poisoning works by hiding malicious instructions inside tool metadata, and because the tool description was never treated as governed data in the first place, the agent reads and acts on it without question. Calling this purely a security issue misses half the picture. It's a data provenance and integrity failure, the same category of failure 5338 already asks organizations to guard against in training data, occurring in a part of the system nobody thought to check.
Read faithfully, 5338's data readiness requirement asks for the full set of MCP servers and tools an agent can reach to be inventoried and checked, not just the data used to train the underlying model. It asks for tool descriptions to go through the same kind of provenance checks training data already goes through. And it asks for something close to a supply chain for tool schemas, built on the same logic 5338 already requires for training data supply chains. None of this is a new category of requirement. It's the same requirement, applied to a part of the system that didn't exist when the requirement was written.
Verification and validation for runtime-defined behavior
5338 calls for verification, validation, and testing of AI systems, and that requirement assumes a system whose behavior can be pinned down at the moment it's assessed. An agent that plans, calls tools, holds state across a session, and produces a multi-step trajectory with real effects on the outside world doesn't sit still long enough for that kind of assessment to fully capture it.
Translating the requirement for agents means extending verification past the model's outputs and into the sequence of tool calls and actions the agent actually takes. And it means runtime guardrails, the execution-time mechanisms that can allow, deny, delay, or escalate an action based on policy, context, identity, or the trajectory so far, become the operational form that verification and validation take once a system's behavior is only fully defined while it's running. Guardrails are what 5338's requirement looks like once it's applied to a system that decides what to do at runtime instead of ahead of time. That reframing matters, because it sets up everything the next two sections depend on: who the agent is when it acts, and how anyone watches what it did.
Rereading identity and access requirements for non-human agents holding live credentials
5338's requirements around operating and maintaining deployed systems carry an implicit demand for access governance. Traditional identity and access management, built around static policies and human login flows, was never designed for agents that can hold credentials across several enterprise systems at once and need their privileges adjusted dynamically, in context, while they're running.
The scale of the problem is easy to state. A single agent might hold credentials for a CRM system, an email account, cloud infrastructure, and a payment system all at the same time. In a multi-agent setup, a lower-privilege agent that gets compromised can use agent-to-agent communication to escalate into a higher-privilege one, a lateral movement pattern that static access policies simply don't account for.
The MCP authorization specification lays out an OAuth 2.1 framework, but it explicitly marks authorization as optional, and MCP server deployments have spread far faster than adoption of that authorization spec has. What the protocol makes possible and what production environments actually enforce are two different things right now.
5338's access and operation requirements call for credentials provisioned just in time, scoped to the specific action an agent is about to take. And they call for identity tied to every single action, through SSO, SCIM, and just-in-time credentials, so an agent's identity can be traced all the way down to the level of an individual tool call.
Microsoft's Entra Agent ID is one example of an organization building toward this. It registers AI agents as their own category inside the identity system, distinct from employee accounts or regular software, with access rules, usage limits, and activity logs assigned per agent. That's not the only shape this kind of response can take, but it shows the identity gap is being treated as something that needs dedicated infrastructure, not a patch on existing human identity systems.
Continuous monitoring across tools, sessions, and organizational boundaries
5338's monitoring and maintenance requirements assume a system with a boundary that holds still. Agents working across MCP servers, outside APIs, and multi-agent pipelines don't have a boundary like that. The edge of the system keeps moving, which turns continuous monitoring into a different engineering problem than the one 5338's language was written to describe.
A full read of 5338's monitoring requirement, applied to agents, has to cover a specific set of threats: tool poisoning and schema poisoning, where malicious instructions sit inside tool metadata and quietly redirect an agent's behavior; prompt injection coming through input that hasn't been sanitized; shadow MCP servers running outside any formal governance process, the agentic version of shadow IT; context getting shared more broadly than it should across sessions multiple agents share; intent drift, where an agent's behavior over the course of a long trajectory slides away from what it was actually authorized to do; and exfiltration carried out through tool calls that look legitimate at every individual step.
Meeting 5338's monitoring requirement for agents, in practice, means detection running at the level of individual tool calls, not just sessions or final outputs. And it means one governance view covering every MCP server, every skill, and every agent in the environment, rather than a patchwork of separate logs nobody can line up against each other.
Retirement and change-control requirements amid continuous tool environment change
Most organizations put real effort into tracking model versions and almost none into tracking tool environment versions, even though for an agent, the tool environment is just as much a part of the system as the model is. 5338 calls for AI systems to be retired responsibly, with a documented reason, and tied into the broader system engineering change-control process. For agents, retirement is a standing obligation, because a change to any connected MCP server, tool schema, or authorization scope is effectively a change to what the agent is capable of doing.
An agent that passed its verification and validation at the moment it was deployed can end up operating outside its original assessed boundary later, simply because a tool provider updated a schema or a marketplace added a new skill the agent now has access to. Without change control that tracks the tool environment and not just the model, the signal that would normally trigger a re-assessment never fires. Nobody notices the system has changed, because on paper, nothing did.
The ClawHub marketplace incident shows what that looks like at scale. Across ClawHub, the marketplace serving the OpenClaw AI agent framework, 1,184 malicious skills were confirmed. Agents that hadn't changed at all were suddenly operating inside a materially different tool environment, with no change-control signal anywhere to flag it.
Meeting 5338's retirement and change-control requirements for agents means version-tracking the full tool inventory, not just the model or the application code sitting around it. And it means having documented criteria, set ahead of time, for when an agent gets suspended or retired because its tool environment has changed enough to invalidate the verification it originally passed. This is the translation work 5338 demands at the end of the lifecycle just as much as at the start, even though the standard's own language never spells it out in these terms.
The process assessment gap in the standards pipeline
The standards world is starting to catch up to all of this, but the infrastructure needed to actually assess whether an organization is doing this translated work, at a measurable level of process capability, doesn't exist yet for agentic systems.
ISO/IEC AWI 25704 is meant to fill part of that gap. Once finished, it will give organizations a real basis for demonstrating process capability. It isn't usable as an assessment tool yet.
On the identity and authorization side, individual contributors have filed active Internet-Drafts with the IETF covering transaction tokens for agents and an Agent Authorization Profile for OAuth 2.0, though neither has been formally adopted by the OAuth Working Group. Neither has been ratified yet.
ISO/IEC 22989, the terminology base 5338 builds on, is being amended to address AI agent definitions directly, and until that amendment finishes, the terminology landscape stays inconsistent from one jurisdiction to the next.
None of this gives enterprises permission to wait. Organizations that do it early are doing what 5338 already asks of them, applied honestly to a system the standard didn't have in front of it when it was written.
Sources
- ISO/IEC 5338:2023 — AI System Lifecycle (International, 2026):
- From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI
- ISO/IEC AWI 25704 - Artificial intelligence — Process assessment — Process assessment model for AI system life cycle processes
- Beyond Component Testing: Validating Agentic AI Systems
- MCP Security Crisis: Systemic Design Flaws in AI Agent Infrastructure


