Field notes · Governing agents

The rulebook for agents is arriving. It asks one thing: show your work

Three rulebooks are landing on AI at once. They come from different places and converge on a single demand: be able to show what your agents did, and why.

Published by Aitonomy We run agent workforces for companies. This is where we write down what that actually takes. What we do

Europe’s AI Act, the ISO management standard your buyers now ask for, and the sovereignty question of who controls your stack are arriving together. They look unrelated. They are not. Each one, in the end, asks you to produce the same thing: a record of what your AI did and on what basis. For agents, which act on their own across many steps, that record is the hard part.

Most teams can get an agent to do the work. Proving afterward what it did, on what basis, and who signed off is where it falls down. And that proof is exactly what the rules are starting to require.

The AI Act: a delay, but not the one people think

The headline in 2026 is that the EU blinked. On 19 November 2025 the Commission published its Digital Omnibus, and after trilogue negotiations the institutions agreed in May 2026 to defer the AI Act’s high-risk obligations. Stand-alone high-risk systems under Annex III move from 2 August 2026 to 2 December 2027, and AI embedded in regulated products moves to 2 August 2028. Law firms tracking it, such as Gibson Dunn, read the deferral as a response to unfinished harmonised standards and national authorities that were not ready.

It would be easy to read that as “compliance can wait”. It cannot. As the AI-law team at Jones Walker put it, August 2 still matters, because most of the Article 50 transparency obligations stayed exactly where they were. From 2 August 2026, Article 50 applies, and it splits its duties by role rather than putting the same one on everybody. Providers must design systems that interact directly with people so that a person knows it is an AI, unless that is already obvious in the circumstances, and providers carry the machine-readable marking of synthetic content. Deployers have their own, narrower duties: emotion recognition and biometric categorisation, deepfakes, and AI-generated text published to inform the public on matters of public interest. That last one has an exception where the text was reviewed by a person and someone holds editorial responsibility for it. Not every AI-generated image, clip or paragraph puts the same obligation on the organisation using it. The obligations for general-purpose models already applied from August 2025. So the burden that lands first is not the heavy high-risk paperwork. It is knowing which of these duties is actually yours, and being able to show what the machine produced and on what basis.

The independent analyst Luiza Jarovsky has tracked every amendment through the Council’s approval in her newsletter, and her point is worth repeating: the postponement is not a pause button.

ISO 42001: the standard your buyers will ask for

The second rulebook is not a law at all. ISO/IEC 42001, published in December 2023, is the first management-system standard for AI, the equivalent of ISO 27001 but for how an organisation governs its AI. It is becoming a procurement filter. Anthropic was one of the first frontier labs to certify, effective January 2025, with AWS and Microsoft following, and Fortune 500 buyers are now asking vendors either to hold the certificate or to show a credible roadmap to it. What the standard actually checks is unglamorous and telling: do you have policies, risk assessments, and monitoring, and can you evidence them. Again, the ask is a record.

Sovereignty: who can be compelled to hand it over

The third rulebook is about control rather than paperwork. A European data centre is not the test.

The legislation in the US doesn’t really care where the data is located. It only cares if it’s operated and controlled by a US entity. Frank Karlitschek, Nextcloud · Holistic Sovereignty podcast

In November 2025 all 27 EU member states signed a declaration to strengthen digital sovereignty and cut strategic dependencies. For anyone running agents over their own and their customers’ code and data, the question of which jurisdiction can compel disclosure is now a first-class design input, not a footnote.

Why this matters for anyone shipping agents

Here is the business line. Agents multiply the number of automated decisions your organisation makes, and every rulebook above raises the price of not being able to account for them. Gartner expects over 40 percent of agentic AI projects to be cancelled by 2027, citing unclear value and inadequate controls. The teams that can show their work will pass audits, win regulated buyers, and keep shipping. The teams that cannot will stall.

Where our platform fits

What all three rulebooks want is not really a legal idea, it is a management one. As Ethan Mollick writes in One Useful Thing, “if you can explain what you need, give effective feedback, and design ways of evaluating work, you are going to be able to work with agents”. Governing an agent is managing it: specify the work, review it, keep a record of both. The rules are making that record mandatory.

This is the layer we build. Aitonomy sits around whichever coding agent you use and turns its activity into something you can govern and evidence.

What the platform provides

One idea · a readable record

  1. Gates.An agent’s change does not count as done until it passes the checks defined for that step. The same gate for a human or an agent pull request.
  2. Approval.A human sign-off is recorded where policy requires one, so autonomy has a named owner rather than an anonymous automated commit.
  3. Audit trail.Every step, spec, check and decision is written down and readable back months later. That is the record the AI Act’s transparency duties, an ISO 42001 audit, and a sovereignty review all end up asking for.

None of that makes an agent compliant on its own. It makes the evidence exist, which is the part teams usually discover they are missing only when someone asks.

The honest edge is worth stating. The omnibus delay is a reprieve, not a repeal, and treating it as cancellation is the mistake that leaves you scrambling in 2027. A certificate is also not the same as being safe: ISO 42001 checks that you have a system, not that the system is good. Governance can curdle into theatre if the record is generated to satisfy an auditor rather than to actually steer the work.

The point is not to produce paperwork. It is to be able to answer, truthfully and quickly, what your agents did and why.

Which of these three is landing on your desk first, and are you set up to answer it?

Common questions

Has the EU AI Act been delayed?+
Parts of it have. The high-risk obligations were deferred in 2026, but other duties stayed on their original timetable, including most Article 50 transparency obligations. The delay creates more preparation time. It does not remove the need to record what an AI system produced and on what basis.
What is ISO 42001, and why does it matter?+
ISO/IEC 42001 is a management-system standard for AI. It asks whether an organisation has policies, risk assessments, monitoring and evidence for how it governs AI. As buyers bring it into procurement, it becomes a commercial question as well as a compliance one.
Does hosting AI in the EU make it sovereign?+
Not on its own. Server location matters, but so does who operates and controls the stack. The practical question is which jurisdiction can compel that organisation to hand over data. A European data centre does not settle that question by itself.
What evidence should you keep about an AI agent’s work?+
Keep enough to reconstruct what the agent did, why it did it, which checks it passed and who approved the result. That record does not make the agent compliant or safe by itself. It is the evidence needed to run the agent in production and gives operators, buyers and auditors something concrete to inspect.

Sources

Mik Nijhuis

Co-founder and Chief AI Officer

Mik has architected enterprise systems for ABN AMRO, Aegon and Shell. He builds agent systems that add throughput without removing the controls complex organisations need.

Aitonomy builds and runs agent roles for recurring business work, supervised in production and measured against a baseline agreed before the first live case.What we do

All insights

Fieldwork

How an agent workforce is actually run. The roles we put into production, and the evidence behind the decisions.

Ready when you are

Bring one workflow.
Leave with a business case.

Bring us the work that slows your teams down. We will map it, count what it costs today, and select the first role only when the numbers support it.