What is AI agent governance, and what does it have to cover?
AI agent governance is the set of decisions an organization makes about software that holds its own credentials: who is accountable for what an agent does, who is allowed to change its behavior, and who writes the standard it is measured against. The identity half is already built.
Questions as you read? Ask, or take the policy template with you.
Where should we send it?
Sent ✓ Download it now →
Microsoft’s Entra Agent ID requires a named human sponsor when an agent is created and transfers that sponsorship to the person’s manager automatically when they leave, and Google issues every agent on its Gemini Enterprise Agent Platform a cryptographic identity of its own.
The half nobody ships is authorship. Microsoft’s own evaluation documentation tells you to confirm that a rubric aligns with “your own judgment” and never says whose judgment that is.
The sponsor is accountable and cannot change the behavior
Somewhere in your company there will be a person whose name sits beside an AI agent, and who is not permitted to change what that agent does.
That is not a forecast. It is written into the product documentation of the largest enterprise software vendor in the world, in two sections of the same page. AI agent governance begins right there, in the distance between the name on the record and the hand on the controls.
Microsoft Entra defines a sponsor as the human “accountable for making decisions about its lifecycle and access.” Further down the same documentation set, sponsors “can’t modify application settings on agent blueprints or agent identities,” and their permissions are limited to what Microsoft calls “nondestructive lifecycle operations.” They can disable an agent. They cannot re-enable it. They cannot change what it says.
So the role carries the consequence and not the control. That single arrangement is the honest starting point, and it is the reason this is an organizational question before it is a technical one.
Nothing here is new, which is the part that should bother you. Microsoft Agent 365 became generally available on 1 May 2026. Microsoft IQ was announced at Build on 2 June 2026. This has been shipping for roughly four months, into tenants that already pay for it, and the accountability question has barely been asked out loud.
The one-page AI use policy.
A 300 word policy you can adopt today, five blanks to fill, and the amnesty that makes it work.
On its way. Check your inbox, or download it now →
What the platforms actually built
The easy version of this argument, that the vendors shipped autonomy and no controls, is false. Here is what exists.
Microsoft treats agent identity as a first-class object in Entra, separate from a human user and separate from a service principal, designed for agents that may be “created and destroyed thousands of times per day.” An agent can optionally be given a user account, which means a mailbox, a OneDrive, Teams presence, and a place in the HR system. It cannot hold a password or a passkey and cannot be assigned privileged administrative roles.
Three human roles attach to it: an owner for technical administration, a sponsor for business accountability, and a manager for the org chart. Managers, Microsoft’s documentation says, “will see agents designated as reporting to them in the Microsoft Entra admin center.”
In a Microsoft session recording made before June 2026, Marco Casalena, a vice president of products for Core AI at Microsoft, demonstrated this and narrated the moment plainly: “when we look in the org chart, we can see that it’s not just an agent. This is my agent. It reports to me in the org chart.”
Google arrived at the same destination from a different direction. Its Gemini Enterprise Agent Platform gives each agent “a strongly attested, cryptographic identity” built on SPIFFE, not shared across workloads, not impersonable, with no long-lived keys. When an agent acts for a person, Google’s documentation says the logs show both the agent’s identity and the user’s.
Anthropic looked at the same problem and declined to settle it in public. Its engineering write-up from 25 May 2026 poses the question rather than answering it: “Should an agent possess its own principal identity, or should it act as an extension of the user and inherit the user’s permissions? Ultimately, the answer may be a blend of the two.” Its containment model is scoped tokens and sandboxing instead of a directory record. For Cowork, “credentials stay in the host keychain, the VM gets a per-session scoped-down token.”
Three named positions, two of them shipping directory-level identity and one of them deliberately holding the question open. This is an industry pattern, not a Microsoft story. If you want the fuller map of how many of these control planes are now competing for the same job, we wrote about AI agent sprawl and why more agents makes a business less intelligent. This piece stays off that ground and asks a narrower question.
One note on the numbers traveling alongside this shift, because they are softer than they sound. The same Microsoft session cited “a recent Gartner study” for an increase in agent accuracy of up to 80% and a cost reduction of up to 60%. The underlying item is a Gartner prediction dated 11 May 2026 from Rita Sallam, a distinguished vice president analyst, and it concerns organizations that prioritize semantics in AI-ready data by 2027. A forward-looking prediction about data modeling is not a completed study about a product layer.
The session also described over a billion agents in operation today; the footnote under Microsoft’s own Agent 365 announcement points to an IDC info snapshot, sponsored by Microsoft, forecasting 1.3 billion AI agents by 2028. A projection is not a census.
The AI Briefing
Tuesdays. 500+ leaders. No hype, just what works.
Where the platform is genuinely enough
Here is the case against this entire article, and it deserves to be made properly before I argue with it.
Agent identity governance is a solved platform problem that consultants are busy reselling as a strategy problem. Entra, Google’s identity model and Anthropic’s sandboxing already cover identity, scope, expiry and audit. Whatever is left over is ordinary management: decide what the agent is for, put someone’s name on it, look at it now and then. A firm that sells governance help has an obvious commercial interest in calling ordinary management a discipline.
I want to concede how much of that is right.
A company of thirty to eighty people, running two or three agents inside one Microsoft 365 tenant, with permissions scoped in Entra and a named sponsor on each agent, has real governance. Not a governance story. Governance. That company does not need us, and anyone telling them otherwise is selling.
Look at what the platform enforces for them. A named human is required when the agent is created and cannot be omitted, which is more than most companies of that size ever managed for their service accounts. Least privilege is the default posture and not an aspiration: agents begin with limited inherited permissions and gain further access through packages that carry expiry dates and approval routing. Attribution is genuinely solved, because actions log against the agent identity and both identities appear when the agent acts for a person.
Most mid-market companies cannot say that about their human systems. And the departure problem is handled by machinery instead of memory.
There is evidence that automated checks beat human gates at this scale, too, and it comes from a vendor arguing against its own convenience. Anthropic’s telemetry, published in May 2026, found that users approved roughly 93% of permission prompts, and that “the more approvals a user sees, the less attention they pay to each, becoming over time much less diligent in their supervision.”
In the same write-up, Claude Code’s auto mode “catches roughly 83% of overeager behaviors before they execute.” Adding a human approval step to a high-volume agent workflow often adds a rubber stamp.
A sixty-person company with two agents, scoped permissions and a named sponsor has real AI agent governance. That is not the part we are arguing about.
That finding deserves a moment of discomfort on our side of the argument, because it sits in tension with something we have published. Our earlier piece on the safety architecture behind trustworthy AI agents leans on human approval as the primary control, drawing on Anthropic’s own five principles. The 93% number undercuts that model at volume. Approval is a strong control when a human still holds the credential and a weak one when they are clicking through a queue. Both things are true, and the second one is newer.
Where agentic AI governance stops being a platform problem
So the disagreement is not about identity, scope or logging. Those are built. The disagreement is about what the platform cannot hold, and there are six places where it stops.
The first is the one I opened with. The accountable person cannot change the behavior. A sponsor is business-accountable for purpose and lifecycle and is explicitly barred from modifying application settings. In every other part of a company we would recognize that arrangement immediately: it is a manager with a budget they are not allowed to spend, or a quality lead who cannot halt the line. We would not call it accountability. We would call it exposure.
The second is that accountability is inherited rather than accepted, which is worth its own section below.
The third is that nothing in the stack says who writes the acceptance criteria, which gets its own section too.
The fourth is that the controls themselves fail quietly. Between January and February 2026, a code defect caused Microsoft 365 Copilot Chat to summarize emails carrying a confidential sensitivity label, from Draft and Sent Items, in spite of data loss prevention policies configured to stop exactly that. The Register reported customer complaints from 21 January 2026 and a fix rolling out through February, under Microsoft reference CW1226324.
Microsoft’s position was that this “did not provide anyone access to information they weren’t already authorized to see,” which is a fair defense and also beside the point. A governance control inside the vendor’s own product did not hold, for weeks, and nothing in the customer’s tenant surfaced it.
The fifth is that the standards bodies have not finished asking the question. The NIST National Cybersecurity Center of Excellence published a draft concept paper in February 2026 called “Accelerating the Adoption of Software and AI Agent Identity and Authorization,” with a comment period that ran from 5 February to 2 April 2026. It is a request for industry input, not guidance.
Among the questions it puts to the industry, verbatim: “How do we ensure non-repudiation for agent actions and binding back to human authorization?” and “How do we establish ‘least privilege’ for an agent, especially when its required actions might not be fully predictable when deployed?”
The OWASP GenAI Security Project released a Top 10 for Agentic Applications on 9 December 2025 with more than a hundred contributors, which tells you the same thing from the security side. The identity layer shipped commercially before the standards work finished defining what it should mean.
The sixth is that the regulation assumes a determinate right answer and does not say who supplies it. Writing in Tech Policy Press on 5 May 2026, Kathrin Gardhouse and Amin Oueslati of The Future Society argue that the EU AI Act, which became generally applicable on 2 August 2026, was not drafted for agents.
Their sharpest point is about accuracy: “The Act’s accuracy metric presupposes a determinate standard against which outputs can be assessed as correct or incorrect, a poor fit for agentic tasks.” Somebody inside the business has to supply that standard. No law names them, and no product creates them.
Accountability that arrives by org chart is not accountability. It is assignment.
Accountability that arrives instead of being accepted
Now the piece of this I find hardest to let go of.
Microsoft’s identity governance documentation describes what happens when a sponsor leaves the company. “Sponsorship of the agent identities is automatically transferred to their manager,” so that, in Microsoft’s words, “there’s always a human user accountable.”
As engineering, that is excellent. It closes the orphaned-service-account problem that has haunted IT departments for thirty years, and it closes it without relying on anybody remembering to do the handover.
As governance, read it again. A manager can arrive on a Monday accountable for an agent whose purpose they never read, whose data reach they never approved, and whose acceptance standard they have never seen. Nobody signed anything. Nothing was explained. The org chart made the decision.
That is precisely the failure mode the AI governance framework we published for companies without a legal team warns about in its third layer, where release authority is drawn by consequence and reversibility, never by seniority. Automatic sponsorship transfer is seniority deciding accountability, dressed as continuity.
The fix is not technical and it is not difficult. It is a five-minute conversation that nobody currently owes anyone. When sponsorship moves, the new sponsor is told what the agent is for, what it can reach, what it has done recently, and what would count as it going wrong. If that conversation cannot happen because nobody can answer those four things, the agent should be disabled until someone can. Which, notably, a sponsor is permitted to do.
Enterprise AI governance programs tend to spend their attention on the moment an agent is created, because that is where the forms are. The moment that actually decides accountability is the moment it changes hands.
The one-page AI use policy.
A 300 word policy you can adopt today, five blanks to fill, and the amnesty that makes it work.
On its way. Check your inbox, or download it now →
Who writes the rule the agent is measured against
This is the gap that no vendor is hiding and no vendor will fill, because it is not theirs to fill.
Every evaluation stack for agents ships with dozens of generic metrics: relevance, groundedness, task completion, tool accuracy. Casalena named the limit of that set exactly in the same session: they measure “does this agent work? Not does this agent work right.” The fix he described, and the one Microsoft’s documentation recommends, is a custom rubric that encodes your business rules.
Read how the documentation hands that over. A rubric, Microsoft writes, “gives you full control over what ‘good’ means for your use case.” It tells you to “iterate on the rubric until it reliably distinguishes between acceptable and unacceptable agent responses,” and to “validate that the rubric scores align with your own judgment before using it at scale.”
Your own judgment. Whose, exactly.
In a company of eight people that question answers itself. In a company of two hundred it does not, because AI model governance at that size means marketing, legal, finance and delivery hold genuinely different views about what an acceptable customer email looks like, and they have never had to reconcile them in writing. The agent does not care. It will apply whichever version got typed into the rubric.
It gets more pointed. When the platform auto-generates a starter rubric, it assigns exactly one criterion a weight between 8 and 10, described as the most outcome-decisive dimension, and everything else a weight between 1 and 6. The generating model chooses which criterion that is.
Microsoft’s own published example encodes operational rules as scored dimensions, giving restaurant booking constraints a weight of 5 and intent recognition a weight of 9. A human may override those weights. Nothing requires one to.
So a policy question, which of our rules matters most when they conflict, gets settled by a number a model picked, inside a file most executives will never open.
The rubric is only hard to write when two people disagree about the rule. No product handles that moment, and no product ever will.
The people who work on evaluation professionally reached this conclusion years before the agent identity layer existed. Hamel Husain, whose evaluation FAQ is one of the more widely read practitioner references in the field, argues for “appointing a single domain expert as a ‘benevolent dictator’” precisely because consensus-built criteria stall, and holds that criteria should emerge from error analysis of real traces instead of being written in advance. He calls the resulting movement criteria drift.
The academic version arrived at the same place: “Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences,” presented by Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran and Arawjo at UIST in 2024, found that evaluation criteria evolve as people look at actual outputs, and that alignment is subjective and iterative, never specifiable up front.
Which means the acceptance standard cannot be procured, delegated to the vendor, or generated. It has to be authored by a person with the standing to be wrong in public, and revised as the organization learns what it actually meant.
Microsoft is building tooling for the mechanical half of this. ASSERT, announced at Build 2026, converts “your policies into concrete, measurable evaluations,” and an Agent Control Specification defines deterministic checkpoints. Both are useful. Neither writes the policy.
What changes in the four layers when the agent holds the credential
Our governance framework for companies without a compliance department asks four questions: what may be used, what may go in, who releases what, and what gets checked. An agent with its own identity does not invalidate that. It splits it.
Layers one and two survive almost intact. Which tools are sanctioned and what data may enter them are the same decisions whether a person or an agent is doing the entering, and the enforcement actually improves, because an agent’s data reach is a permission set you can inspect, where a person’s habits can only be trusted. If you want the artifact that puts those two layers in front of your team, the AI acceptable use policy we published includes a complete one-page version to copy.
Layers three and four are where the model breaks, and it breaks for one reason. Both were designed around a person who can be asked why.
Layer three, who releases what, assumes a reviewer who can interrogate the work. You can ask a colleague what they were thinking, hear an answer, and decide whether to trust the next thing they produce. An agent will produce an explanation on request, fluently, and that explanation is generated rather than recalled. The release decision has to move upstream, from reviewing output to authorizing a class of action against a written standard. This is the point where AI oversight stops meaning supervision and starts meaning specification.
Layer four, what gets checked, assumes a review cadence that human work volume made possible. An agent operating continuously produces more output in a week than your review list was ever sized for, which is why sampling against an explicit rubric replaces reading, and why the rubric becomes the governance artifact. The audit log tells you what happened. Only the rubric tells you whether it was right.
There is a third thing that arrives with all of this, and it belongs to architecture more than to governance. Agents call other agents. A context service is itself a sub-agent. The moment one agent’s output becomes another agent’s input, you are no longer governing a tool, you are governing a chain, which is the territory where loops become graphs and orchestration turns into a design discipline of its own.
What should an AI governance checklist for agents cover?
Not a policy template. We already published one of those, and a template cannot contain the decision this article is about.
What follows is the short list of things that have to be true before an agent with its own credentials should be allowed to act on your behalf. Five items. If you cannot complete them for an agent you are already running, you have found your work for this week.
The fifth item is the whole argument compressed. Every organization has rules that have never been reconciled because they have never had to be applied by the same hand on the same day. An agent applies them all, at speed, without the social skill of noticing that two of them conflict.
That is where the leftover stops being ordinary management. It is not that naming an owner is hard. It is that the agent forces a decision the organization has been quietly deferring for years, and it forces it in writing.
Where an outside partner actually helps, and where they do not
I will be specific about this, because the honest version is narrower than a consulting firm would like.
You do not need help naming a sponsor, scoping permissions, or turning on identity governance. Your IT partner or your own administrator can do all of it, and the documentation is good.
What is hard from inside is the fifth checklist item. Getting two departments to agree, in writing, on what correct output looks like when their current answers differ is a political act, not a technical one, and the person inside the company who calls that meeting has to keep working with everyone in it afterward.
That is the moment an outsider is genuinely useful: someone who can hold the disagreement open, write down what was actually decided, and leave. The relevant work on our side is AI strategy, and it is a short engagement or it is not working.
Apply one test to us or to any firm you consider. Ask what you will own when they leave. If the answer is a document with your names in it, a rubric your own people wrote and can revise, and a named route for the next disagreement, that is ownership transfer. If the answer is a dashboard only they can read or a framework that requires their presence to run, you are renting judgment, which is the one thing you cannot afford to rent.
If you want a starting point that costs nothing, our AI use policy guide is the artifact most teams are missing before any of this becomes relevant.
Write The Accountability Record For One Agent
Pick the agent in your company with the most reach and the least documentation. Paste this into your AI assistant. It interviews you until the five checklist items are answered, and returns the record you can hand to a sponsor.
Context: I am writing an accountability record for one AI agent running inside my company. It is a [DESCRIBE THE AGENT] and it currently has access to [LIST SYSTEMS AND DATA]. Interview me one question at a time. Push back when my answer is a job title instead of a person, and when I describe a process we do not actually run.
Step 1. The sponsor: Ask who is accountable for this agent by name. Then ask whether that person has explicitly agreed, in a conversation, and when. If they have not, record that as the first gap.
Step 2. The purpose: Make me state what this agent is for in one sentence a client could read. If I need more than one sentence, keep asking until the scope narrows or I admit it has not been decided.
Step 3. Consequence and control: Ask who is accountable and who can actually change the agent’s behavior or instructions. If those are different people, ask how a change request travels between them and how long it takes in practice.
Step 4. The acceptance standard: Interview me until I can describe what good output looks like for this agent, with two examples of output that would be unacceptable and why. Do not write the standard for me. Draft only what I have said.
Step 5. The disagreement route: Ask what happens when two departments judge the same output differently. Who decides, in what forum, by when. Record UNDECIDED instead of filling it in.
Output: A one-page accountability record with the five answers, an UNDECIDED list kept separate at the end, and a single recommended next action.
The UNDECIDED list is the useful part. Those are the questions your company has never settled, which is why no platform setting and no policy template could have answered them for you. See where you stand →
The stand
The agent got a badge, a mailbox and a line in the org chart before anyone asked what it means for a person to answer for it.
That was not negligence on the vendors’ part. Microsoft, Google and Anthropic each built what they are actually able to build: identity, scope, expiry, logs, and in Anthropic’s case an honest public admission that the principal question is unresolved. None of them can build the part that matters most, because the part that matters most is a judgment your company has to make and then defend.
Humans First has never been a comfort. It is a description of where the load sits. When the thing being governed holds its own credentials, the human role does not shrink to supervision. It concentrates into the two acts no system will perform for you: accepting accountability on purpose rather than inheriting it, and writing down what right looks like when your own people disagree.
Every AI agent running in your company right now is executing somebody’s definition of good. Find out whose.
Sources
- What are agent identities? (Microsoft Learn): agent identity as a first-class Entra object, distinct from users and service principals
- Agent owners, sponsors, and managers (Microsoft Learn): the three human roles, sponsor required at creation, and the restriction on modifying application settings
- Agent users (Microsoft Learn): the optional agent user account, mailbox, OneDrive and Teams presence, and its limits
- Agent ID governance overview (Microsoft Learn): automatic sponsorship transfer to the sponsor’s manager, access packages, licensing requirements
- Microsoft Agent 365 overview (Microsoft Learn): scope, licensing and general availability
- Microsoft Agent 365 now generally available (Microsoft Security Blog, 1 May 2026): GA date; also the source of the IDC 1.3 billion agents by 2028 footnote
- Rubric evaluators (Microsoft Learn): “full control over what ‘good’ means,” rubric auto-generation, and the 8 to 10 weighting of a single decisive criterion
- What’s new in Microsoft Foundry at Build 2026 (Microsoft DevBlogs, June 2026): Microsoft IQ announcement date, ASSERT, Agent Control Specification
- Agent identity overview (Google Cloud): SPIFFE-based cryptographic agent identity and dual-identity logging
- How we contain Claude (Anthropic, 25 May 2026): the open question on principal identity, the 93% permission-approval telemetry, and the 83% auto-mode catch rate
- Accelerating the Adoption of Software and AI Agent Identity and Authorization (NIST NCCoE draft concept paper, February 2026): open questions on non-repudiation, human binding and least privilege
- OWASP Top 10 for Agentic Applications 2026 (OWASP GenAI Security Project, 9 December 2025): agentic risk taxonomy, 100+ contributors
- The EU AI Act Is Not Ready for Agents (Tech Policy Press, 5 May 2026): Gardhouse and Oueslati on the accuracy metric and its assumption of a determinate standard
- Ahead of the Curve: Governing AI Agents under the EU AI Act (The Future Society, 4 June 2025): how agents fall under existing GPAI and high-risk provisions
- Microsoft Copilot summarized confidential emails despite DLP policies (The Register, 18 February 2026): reference CW1226324, customer reports from 21 January 2026
- Primera notificación de brecha de datos personales causada por un ataque ejecutado mediante un agente de IA (AEPD, 14 September 2026): first agentic breach notification received by the Spanish regulator
- Evals FAQ (Hamel Husain, updated 18 September 2026): the benevolent dictator argument and criteria drift
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences (Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran, Arawjo, UIST 2024): evaluation criteria evolve as people examine real outputs
- Gartner: lack of semantics causes inaccurate AI agents and wasted spending (Fortune, 19 May 2026, reporting Gartner’s 11 May 2026 prediction by Rita Sallam): the correct framing of the 80% accuracy and 60% cost figures as a 2027 prediction about semantics in AI-ready data
- Microsoft discovery session recording featuring Marco Casalena, vice president of products for Core AI, Microsoft. Recorded before June 2026; no public URL.
Frequently Asked Questions
What is AI agent governance?
AI agent governance is the set of decisions an organization makes about AI agents that hold their own credentials and act without a person in the loop for every step. It covers who is accountable for what the agent does, who may change its behavior, what data and systems it may reach, and who authors the standard its output is measured against. The identity and permission parts are increasingly handled by the platform. The accountability and acceptance-standard parts are organizational decisions no product makes for you.
How is agent governance different from AI governance generally?
General AI governance answers four questions about people using AI tools: which tools may be used, what data may go into them, who releases AI-assisted work, and what gets checked afterward. The first two survive intact when an agent holds its own identity. The last two break, because both were designed around a person who can be asked why they made a choice.
An agent produces an explanation on request, but that explanation is produced on demand and not recalled from experience, so the release decision has to move from reviewing individual output to authorizing classes of action against a written standard.
Who is accountable when an AI agent makes a mistake?
Inside Microsoft’s model, the sponsor is, by definition. Entra documentation describes sponsors as human users “accountable for making decisions about its lifecycle and access,” and a sponsor is required when an agent identity is created. The complication is that the same documentation bars sponsors from modifying application settings on agent blueprints or identities, so the accountable person often cannot change the behavior they are accountable for. Legally and practically, accountability lands on the organization, which makes it worth deciding internally who holds it and giving that person a route to act.
Do AI agents need their own identities?
Two of the three major vendors say yes and have shipped it. Microsoft treats agent identity as a first-class object in Entra, distinct from a user or a service principal. Google issues each agent a cryptographic identity built on SPIFFE, with logs showing both agent and user when the agent acts on someone’s behalf. Anthropic has publicly left the question open, asking whether an agent should possess its own principal identity or inherit the user’s permissions, and answering that the resolution may be a blend of both.
The practical benefit of a separate identity is attribution: you can tell what the agent did as opposed to what the person did.
What should an AI agent governance policy cover?
Five things, and a policy template will only carry the first two. Name a sponsor who has explicitly agreed and not merely been assigned. State the agent’s purpose in one sentence. Write down separately who is accountable and who can change the behavior, plus the route between them if they differ. Have one named person author the acceptance standard for the agent’s output, with examples of unacceptable output. And decide in advance what happens when two departments judge the same output differently, because that is the item everyone skips and the only one that gets tested.
Is Entra Agent ID or Agent 365 enough on its own?
For many small and mid-sized companies, closer to enough than they expect. Agent 365 reached general availability on 1 May 2026, and the surrounding identity governance requires either Microsoft 365 E7 or Agent 365 with at least Entra P1, so it is not free with a standard tenant.
What it genuinely enforces is real: a named human at creation, least privilege by default, access packages with expiry and approval routing, and attribution in the logs. What it does not do is decide what the agent is for or what correct output looks like. It also is not infallible, as reference CW1226324 showed when a defect caused Copilot Chat to summarize labeled confidential email in spite of configured data loss prevention policies.
Do regulators treat AI agents differently yet?
They are starting to. In September 2026 the Spanish data protection authority, the AEPD, published the first personal data breach notification it has received attributable to an attack carried out by an AI agent, describing the shift from an AI assisting an attacker to the AI acting agentically and noting that this increases speed, scale and adaptability. That case involved an attacker’s agent and not a company’s own, so it is not evidence of an internal governance failure. It is evidence that a regulator now treats agent autonomy as its own category.
Separately, analysts at The Future Society argue the EU AI Act, generally applicable since 2 August 2026, was not drafted for agents, particularly its accuracy requirement, which assumes a determinate standard of correctness that somebody inside the business has to supply.



