Whether it agrees with your practitioners.
Held-out evaluation measures behaviour on evaluation data. Agreement on live matters requires a recorded disposition history, which is a separate exercise.
MKC2 results ↗Legal AI evaluation is mostly the work of writing down what counts as defective. The measurement is straightforward once that exists.
01 / DEFINITION
The short answer first.Legal AI evaluation is the measurement of an AI system's performance on legal work against an explicit standard for what counts as defective. It requires a held-out set drawn from real material, checks defined by practitioners, and reporting per check rather than as a single accuracy figure.
The reason the standard comes first is that legal quality is multidimensional in a way general evaluation is not. A passage can be well drafted and wrong on the law, correct and unsupported by its cited authority, accurate and procedurally inappropriate for the matter. A single quality score collapses those into a number nobody can act on.
02 / BUILDING THE STANDARD
Decomposition into checkable questions.A workable standard is a bank of specific questions, each with a defined answer space and each narrow enough that two practitioners agree on the answer. MKC2's framework applies 45 such checks to every item. The decomposition is the expensive part and it is also the durable part: it survives model changes, which a prompt does not.
Three properties make a check usable:
03 / THE WORKED EXAMPLE
MKC2, reported in full.Privilege AI's calibration record for MKC2, dated 21 September 2026, measures 45 checks on 4,403 held-out items, with both the model being replaced and the stock Qwen3-8B base re-scored in the same run on the same material.
| Measure | MKC2 | Stock base | Model replaced |
|---|---|---|---|
| Ranking quality | 0.9623 | 0.7281 | 0.6457 |
| Answers correct | 93.17% | 83.09% | 80.73% |
| Brier score | 0.0620 | 0.1274 | 0.1409 |
| Calibration error | 0.0610 | 0.1363 | 0.1785 |
Calibration record, 21 September 2026. Benchmark results on held-out evaluation data, not live-matter outcomes. No attorney dispositions have been recorded, so none of these figures show that MKC2 agrees with a lawyer on live work. One check of 45 is refused outright for insufficient ranking quality; 25 carry report-only authority. MKC2 reports; it does not act. Full record on the MKC2 page.
The ranking figure is the one that decides whether the system is usable. A score that orders defective work above sound work only loosely cannot support a review queue, whatever its accuracy: no threshold separates the classes. The model being replaced scored 0.6457 there, below the stock base model, which is the clearest statement of why the work was done.
04 / WHAT A BENCHMARK CANNOT TELL YOU
The limits a practice should insist on hearing.Held-out evaluation measures behaviour on evaluation data. Agreement on live matters requires a recorded disposition history, which is a separate exercise.
Performance is corpus-dependent. Vocabulary, document types and the definition of a defect vary between practices.
A system that finds defects nobody has time to read has not improved anything. The review workflow is part of the measurement.
Two questions are worth asking any supplier. Were the baselines re-scored in the same run, or quoted from elsewhere? And which checks failed to reach the bar? A report with no negative findings has either an unusually good model or an unusually incurious author. See AI model benchmarking and held-out evaluation.
05 / QUESTIONS
Asked when assessing a legal AI system.Accuracy alone is the wrong question. Ask for ranking quality — whether the system orders defective work above sound work — and for calibration, because those decide whether the output can be used to direct review. A high accuracy figure with weak ranking quality cannot support a queue.
Enough for every reported check to be measurable, which for a decomposed standard means thousands rather than hundreds. MKC2's record uses 4,403 items precisely so that individual checks can be reported separately rather than averaged.
The practitioners who carry the consequence, not the engineering team. A standard written for measurement convenience drifts towards what is easy to score and loses the confidence of the people expected to rely on it.
As a filter, not as a decision. Performance is corpus-dependent, and a figure produced on someone else's documents with someone else's defect definition says little about your material. Run your own held-out set.
That a check records a finding for a human reader and changes nothing on its own. It is the tier assigned where a check's measured evidence supports reporting but not gating. Twenty-five of MKC2's 45 checks sit there by design.
Privilege AI builds held-out evaluation sets and per-check reporting, and publishes the checks that do not reach the bar.