Nothing changes after approval.
Deployment is governed once and then never revisited, while the model, the prompt, the index and the input distribution all move. Approvals should expire.
MKC2 results ↗Most AI governance documents describe a system nobody operates. The version that holds up is an inventory, a set of owners, and measurement attached to each decision.
01 / DEFINITION
The short answer first.AI governance is the set of processes by which an organisation decides what AI systems it will operate, what each may do, who is accountable for it, and what evidence is required before and after deployment. Its output is a record, not a document.
The distinction between a policy and a governance system is whether it can answer specific questions about specific systems. What models are in use, on what data, approved by whom, on what evidence, with what measured performance, and what happens when one fails. A policy states principles; a governance system answers those six questions on demand.
02 / THE INVENTORY PROBLEM
Governance starts with knowing what exists.Nearly every AI governance effort discovers the same thing first: the organisation does not know what it is running. Models arrive inside SaaS features, in scripts written by analysts, in browser extensions and in pilot projects that quietly became load-bearing. A register built by asking teams to self-declare captures the systems whose owners considered themselves in scope.
Inventories that stay accurate are attached to something teams need: a credential, a network route, a model gateway, a procurement step. The control creates the record as a side effect, which is the only mechanism that survives contact with a busy organisation.
| Field | Question it answers |
|---|---|
| Purpose and use case | What is this system for, and who relies on it? |
| Model and version | What exactly is running, and can it be reproduced? |
| Data reached | What could this system disclose if it misbehaved? |
| Deployment location | Does anything leave our boundary? |
| Authority level | Can it act, or only report? |
| Evaluation evidence | What was measured, when, and on what data? |
| Accountable owner | Who answers for it? |
| Withdrawal procedure | How is it switched off, and has that been tested? |
03 / EVIDENCE BY RISK LEVEL
Proportionality, made concrete.Requiring a full evaluation record for a system that drafts internal meeting notes produces a governance process people avoid.
A workable scheme ties evidence to consequence. Systems whose output is discarded or trivially checked need registration and an owner. Systems whose output informs a professional judgement need a held-out evaluation with per-check reporting. Systems that can act need that, plus trajectory evaluation, plus a demonstrated reversal path.
The evidence itself is what an evaluation framework produces: defined checks, held-out data, per-check thresholds, published refusals and a dated record. Privilege AI's own published work is structured that way — 45 checks scored on 4,403 held-out items, with one check refused outright and 25 held at report-only authority, and an explicit statement that the figures are benchmark results rather than live-matter outcomes. See AI evaluation framework and the MKC2 record.
04 / WHAT GOVERNANCE OFTEN MISSES
Three gaps that recur.Deployment is governed once and then never revisited, while the model, the prompt, the index and the input distribution all move. Approvals should expire.
A third-party feature that reads client data is an AI deployment whoever built it, and its model can change without notice.
Oversight is claimed as a control with no measurement of whether reviewers have the capacity to exercise it. See human-in-the-loop AI.
The first gap is the most consequential, because it is invisible: nothing breaks, the register still lists an approved system, and the thing actually running bears a decreasing resemblance to the thing that was approved. Attaching an expiry to every approval, and requiring current measurement to renew it, is the cheapest available fix. The measurement side is covered in AI model monitoring.
05 / QUESTIONS
Asked when AI governance is being set up.Governance decides what is allowed and who is accountable. Risk management identifies what could go wrong, how likely and how badly, and what is being done about it. Governance consumes risk assessments to make decisions. See AI risk management.
Some central function is needed to hold the inventory and the standards. A committee that must approve every use case becomes a bottleneck and is bypassed. The pattern that works is central standards with distributed, named accountability.
Inventory it like any other AI system, and ask two extra questions: what data leaves the boundary, and what notice you get when the model changes. The second is frequently none, which is itself a finding worth recording.
Proportionate to consequence: an owner and a registration at minimum; a held-out evaluation with per-check results where output informs a professional judgement; trajectory evaluation and a tested reversal path where the system can act.
Badly designed governance does, by making approval a negotiation. Governance that specifies exactly what evidence is required for each risk level speeds deployment up, because teams know the target before they start.
Privilege AI produces the evidence governance decisions rest on: held-out evaluation, per-check authority and published refusals.