Compliance theatre starts where evidence stops
Theatre is not caused by people being dishonest. It is what a control turns into when it stops producing an artifact, and you can find every instance of it in your own register in an afternoon.
The phrase “compliance theatre” is normally used as an accusation, aimed at people who are assumed to be either lazy or cynical. That framing has kept the industry from noticing something more useful: theatre has a precise boundary, it is structural rather than moral, and you can locate it in your own control set without interviewing anybody.
My position: a control is theatre from the moment its operation stops producing an artifact that someone who was not there could check later. Nothing about intent enters into it. Diligent people running a control faithfully every single time are still producing theatre if the running of it leaves no trace, because a control nobody can verify is indistinguishable from one nobody ran.
Declare the obvious interest: this is what my employer sells into. Discount me accordingly, and then run the test yourself, because the test costs nothing and the answer belongs to you.
The boundary, stated exactly
Take any control you claim. “All changes to production are reviewed by a second engineer.” “Agent-initiated changes require an approved ticket.” “Access to customer data is logged.”
Ask four questions in order.
Most control registers I have seen fail at question one for somewhere between a third and half of their entries. Not because anyone is faking. Because the controls were written as descriptions of behaviour, and behaviour does not leave artifacts unless the system it runs through is built to capture them.
Why agentic delivery moved the line
This was survivable when the volume was low, for a reason nobody likes to say out loud: human memory was an acceptable fallback. When a control had no artifact, you could still reconstruct what happened by asking the four people involved, and the reconstruction was usually good enough for an internal audit.
That fallback is going. Not because people are leaving faster, but because the number of acts per person went up and the fraction of them any individual can recall went down. Ask an engineer what happened in a specific change from four months ago and, when a large share of the work was initiated on their behalf and merged under a service account, the honest answer is that they do not know.
So controls that were quietly relying on memory are now relying on nothing, and the register still says they are operating.
The specific standard now on the table
The EU AI Act’s enforcement phase began on 2 August 2026, and Article 12 is worth reading in the original because it is unusually concrete about this exact boundary. It requires automatic logging of events relevant to risk and traceability, tamper-evident, retained for six months, and twenty-four for biometric and law-enforcement uses. There are separate transparency obligations: AI systems disclosing themselves, and generated content being marked.
Read that as a definition rather than as an obligation and it maps almost exactly onto the four questions above: automatic covers question two, tamper-evident covers question three, and the retention period covers question four. Whether or not the Act applies to your systems, a regulator has now written down what it thinks counts as evidence, and that definition will travel.
One thing I will not say, and will not let a vendor say on my behalf: no tool makes an organisation compliant with the AI Act or anything else. Tools produce or fail to produce evidence. Compliance is a judgement made by people about your whole operation, and evidence is what that judgement is made from.
What theatre actually costs you
Not the audit finding. That is the cheap outcome.
The expensive one is that theatre consumes the budget which would otherwise have fixed the gap. A control that is written down, reported as green, and reviewed quarterly looks handled. It occupies the slot. Nobody funds a project to instrument something that is already marked as operating, so the theatre is not merely useless, it is actively load-bearing against the fix.
Which produces an uncomfortable conclusion I have come to believe: an admitted gap is worth more than an unevidenced control. The gap gets a remediation plan. The unevidenced control gets a tick.
Control as claim
- "Changes are reviewed before merge"
- Evidence: the policy document
- Verified by: asking the team lead
- Fails silently, indefinitely
- Reported green all the way down
Control as record
- "Each merge carries an approval event with an identity and a timestamp"
- Evidence: the event itself
- Verified by: querying it
- Fails visibly, as a gap in the data
- Reported by counting, not by asking
Where this breaks down
Plenty of real controls legitimately leave no artifact. Design judgement, a well-run threat modelling session, an experienced engineer deciding a change smells wrong. These are among the most effective controls any organisation has, and my test marks them all as theatre. That is a genuine limitation of the test, not a hidden strength.
Evidence can be produced without the control operating. Once you make an artifact the target, the artifact becomes the thing people satisfy. An approval event with an identity and a timestamp is trivially produced without reading anything. I have described a necessary condition for a real control and dressed it slightly as a sufficient one.
This can become its own theatre, more expensively. Instrumenting everything produces enormous volumes of records that nobody queries, which is theatre with a storage bill. The organisations I have seen do this well instrument a small number of controls thoroughly, and the ones doing it badly instrument everything shallowly.
“An admitted gap beats an unevidenced control” is not universally safe advice. In some regulated settings, formally recording a gap has consequences that a careful person should take seriously before acting on a paragraph in an article. Take that one to your own counsel rather than to me.
And the Act’s definition may not survive contact with practice. Article 12 is new, enforcement is early, and how supervisory authorities actually read “tamper-evident” is not yet established. I am treating a young definition as a stable one because it is the best available, not because it is settled.
The takeaway
Theatre is not a character flaw. It is the state a control decays into once running it stops producing something checkable, and for a long time human memory hid the decay.
The volume of agentic delivery has removed that cover, and a regulator has now put a fairly precise definition of evidence on the table. Both point the same way: judge your controls by what they emit, not by what they assert.
If you take one thing into next week: pick the five controls you would most hate to be wrong about, and for each one find the artifact its last operation produced. Not the policy. The artifact. The ones where you cannot are your real register.