Every AI governance framework asks for evidence. ISO/IEC 42001 expects documented information. The EU AI Act expects technical documentation, logs and records. The NIST AI RMF expects measurement. Internal audit expects something it can test. Almost none of them say what a piece of evidence has to look like to be worth anything.
That gap is where most AI governance quietly fails. Not because organisations refuse to produce evidence, but because what gets produced is a claim wearing the clothes of an artefact.
A claim is not evidence
Consider the sentence testing completed: yes, ticked in a governance form. It is not evidence. It is an assertion about evidence that may or may not exist, made by someone who may or may not have seen it, about a version of the system that may or may not be the one now in production.
The same applies to a great many things routinely filed as evidence: a policy document offered as proof that a control operates, a screenshot with no version attached, a vendor's marketing page describing a safety feature, a model card written by the model's own developer, an architecture diagram that predates two releases.
Five properties of evidence that holds up
Usable governance evidence has five characteristics. They are unglamorous, and they are what separates a governance record from a governance folder.
1. It exists as an artefact
There is a document, a test output, a log extract, a signed conclusion — something with content, that can be opened by a person who was not in the room.
2. It is about this system, at this version
An evaluation of the model as it was two releases ago is evidence about that model, not this one. This is the property most often lost, and the easiest to lose: a system changes, the evidence stays where it was, and the link between them silently stops being true. Evidence that is not bound to a version has a shelf life nobody is tracking.
3. It is dated and attributable
Someone produced it, on a date, and can be named. Anonymous evidence cannot be questioned, and evidence that cannot be questioned cannot be relied on.
4. It addresses the specific risk it is offered against
This is the subtle one. Aggregate accuracy is evidence about aggregate accuracy. Offered against a fairness concern, it is not evidence at all — a model can improve on an overall metric while getting materially worse for a subgroup, and the aggregate figure will conceal it. Evidence must be pointed at the thing you are worried about, not at the thing that was easy to measure.
5. It can be retrieved in a year
Evidence in a departed employee's mailbox, in a chat thread, or in a personal drive is evidence you do not have. The question a governance record has to survive is not did we look at this? but can we show, eighteen months later, what we looked at and what we concluded?
What good looks like, by area
For most AI use cases, a proportionate evidence set is smaller than people fear. What matters is that each item is real.
- Inventory entry — the system exists in a list, with an owner, a purpose and a version.
- Intended use and limits — what it is for, and the uses explicitly considered out of scope.
- Data provenance and lawful basis — where the data came from, what may be done with it, and what the provider may do with it.
- Evaluation of the failure modes that matter — not only overall performance, but the specific ways this system could plausibly go wrong, tested on data that includes the hard cases.
- Human oversight procedure — who reviews the output, when they must intervene, what happens when the system is wrong, and how it is switched off. A procedure that cannot answer the third question is not an oversight arrangement.
- Monitoring and thresholds — what is watched in production, at what level someone acts, and who receives the alert.
- Rollback or stop route — the named person and the technical means to interrupt the use.
- Third-party assurance — what the provider has actually evidenced, for which configuration, on whose population.
- The decision — who approved the use, on what basis, with what conditions, against which version.
Weak evidence, and why it passes
Certain artefacts are accepted far more readily than they deserve, usually because they arrive looking official.
A vendor's bias audit. Frequently genuine, and frequently not about you: run by the vendor, on the vendor's population, for a configuration that is not yours, at a date that is not recent. It is evidence that the vendor tested something. Whether it is evidence about your deployment is a separate question, and one that should be answered in writing by whoever accepts it.
Aggregate performance figures. Discussed above. A single headline metric is the most common way a subgroup harm reaches production undetected.
The policy. A policy states an intention. Evidence shows the intention was carried out. The two are filed in the same folder and are not the same thing.
“A human is in the loop.” Almost always stated, rarely described. If nobody can say what the reviewer is expected to notice, how much time they have, and what they do when they disagree with the system, the loop is decorative.
Absence of evidence is a finding, not a failure
It is worth being clear about what missing evidence means. It does not automatically mean the AI use is unsafe or must be stopped. Plenty of low-impact uses do not warrant a formal evaluation, and saying so explicitly is a legitimate governance outcome.
What matters is that the absence was noticed and decided by someone accountable, before deployment, rather than discovered afterwards by someone investigating an incident. A record that says “no fairness evaluation was performed; the governance lead accepted this on the basis that the system does not affect individual outcomes” is a governed position. Silence is not.
Collect it at the moment of the decision
The practical difficulty with evidence is not knowing what is needed. It is that evidence is easy to attach when the decision is being made and extremely expensive to reconstruct a year later, when the people have moved on and the reasoning lives in nobody's memory.
Organisations that do this well share one habit: the evidence is requested at the point the change is proposed, by the process that approves it, and the request names the specific artefact rather than asking whether testing was done. It is a smaller intervention than it sounds, and it is the difference between a governance function that can answer questions and one that can only ask them.
What is AI governance — and why does AI need more governance than traditional software?
Traditional software controls still matter. AI adds characteristics that ordinary software governance was not designed to answer on its own.
Read the first insightPrimary sources
- ISO/IEC 42001:2023 — Artificial intelligence management system
- European Union — Regulation (EU) 2024/1689 (Artificial Intelligence Act)
- NIST — Artificial Intelligence Risk Management Framework
- NIST — How AI Risks Differ from Traditional Software Risks
- ICO — Guidance on AI and data protection
This article is general information about AI governance. It is not legal advice, certification guidance or a statement that every cited requirement applies to every AI system. Regulatory timetables in this area move; check the current position before relying on it.
