Privilege AIMKC2 results

Policy is cheap.
Evidence is not.

Most AI governance documents describe a system nobody operates. The version that holds up is an inventory, a set of owners, and measurement attached to each decision.

READ THE DEFINITION ↓INVENTORY · OWNERS · EVIDENCE

01 / DEFINITION

The short answer first.

What is AI governance?

AI governance is the set of processes by which an organisation decides what AI systems it will operate, what each may do, who is accountable for it, and what evidence is required before and after deployment. Its output is a record, not a document.

The distinction between a policy and a governance system is whether it can answer specific questions about specific systems. What models are in use, on what data, approved by whom, on what evidence, with what measured performance, and what happens when one fails. A policy states principles; a governance system answers those six questions on demand.

02 / THE INVENTORY PROBLEM

Governance starts with knowing what exists.

Nearly every AI governance effort discovers the same thing first: the organisation does not know what it is running. Models arrive inside SaaS features, in scripts written by analysts, in browser extensions and in pilot projects that quietly became load-bearing. A register built by asking teams to self-declare captures the systems whose owners considered themselves in scope.

Inventories that stay accurate are attached to something teams need: a credential, a network route, a model gateway, a procurement step. The control creates the record as a side effect, which is the only mechanism that survives contact with a busy organisation.

What an inventory entry has to hold for governance to be possible, and the question it lets you answer.
FieldQuestion it answers
Purpose and use caseWhat is this system for, and who relies on it?
Model and versionWhat exactly is running, and can it be reproduced?
Data reachedWhat could this system disclose if it misbehaved?
Deployment locationDoes anything leave our boundary?
Authority levelCan it act, or only report?
Evaluation evidenceWhat was measured, when, and on what data?
Accountable ownerWho answers for it?
Withdrawal procedureHow is it switched off, and has that been tested?

03 / EVIDENCE BY RISK LEVEL

Proportionality, made concrete.

Ask for proof
in proportion.

Requiring a full evaluation record for a system that drafts internal meeting notes produces a governance process people avoid.

A workable scheme ties evidence to consequence. Systems whose output is discarded or trivially checked need registration and an owner. Systems whose output informs a professional judgement need a held-out evaluation with per-check reporting. Systems that can act need that, plus trajectory evaluation, plus a demonstrated reversal path.

The evidence itself is what an evaluation framework produces: defined checks, held-out data, per-check thresholds, published refusals and a dated record. Privilege AI's own published work is structured that way — 45 checks scored on 4,403 held-out items, with one check refused outright and 25 held at report-only authority, and an explicit statement that the figures are benchmark results rather than live-matter outcomes. See AI evaluation framework and the MKC2 record.

04 / WHAT GOVERNANCE OFTEN MISSES

Three gaps that recur.
↳ 01

Nothing changes after approval.

Deployment is governed once and then never revisited, while the model, the prompt, the index and the input distribution all move. Approvals should expire.

↳ 02

Vendor systems are out of scope.

A third-party feature that reads client data is an AI deployment whoever built it, and its model can change without notice.

↳ 03

Human review is assumed, not measured.

Oversight is claimed as a control with no measurement of whether reviewers have the capacity to exercise it. See human-in-the-loop AI.

The first gap is the most consequential, because it is invisible: nothing breaks, the register still lists an approved system, and the thing actually running bears a decreasing resemblance to the thing that was approved. Attaching an expiry to every approval, and requiring current measurement to renew it, is the cheapest available fix. The measurement side is covered in AI model monitoring.

05 / QUESTIONS

Asked when AI governance is being set up.

What is the difference between AI governance and AI risk management?

+

Governance decides what is allowed and who is accountable. Risk management identifies what could go wrong, how likely and how badly, and what is being done about it. Governance consumes risk assessments to make decisions. See AI risk management.

Do we need an AI governance committee?

+

Some central function is needed to hold the inventory and the standards. A committee that must approve every use case becomes a bottleneck and is bypassed. The pattern that works is central standards with distributed, named accountability.

How does governance handle AI inside third-party software?

+

Inventory it like any other AI system, and ask two extra questions: what data leaves the boundary, and what notice you get when the model changes. The second is frequently none, which is itself a finding worth recording.

What evidence should governance require before deployment?

+

Proportionate to consequence: an owner and a registration at minimum; a held-out evaluation with per-check results where output informs a professional judgement; trajectory evaluation and a tested reversal path where the system can act.

Does governance slow AI deployment down?

+

Badly designed governance does, by making approval a negotiation. Governance that specifies exactly what evidence is required for each risk level speeds deployment up, because teams know the target before they start.

Governance
needs measurement.

Privilege AI produces the evidence governance decisions rest on: held-out evaluation, per-check authority and published refusals.