Privilege AIMKC2 results

Define the defect.
Then measure.

Legal AI evaluation is mostly the work of writing down what counts as defective. The measurement is straightforward once that exists.

READ THE DEFINITION ↓STANDARD · HELD-OUT · PER-CHECK

01 / DEFINITION

The short answer first.

What is legal AI evaluation?

Legal AI evaluation is the measurement of an AI system's performance on legal work against an explicit standard for what counts as defective. It requires a held-out set drawn from real material, checks defined by practitioners, and reporting per check rather than as a single accuracy figure.

The reason the standard comes first is that legal quality is multidimensional in a way general evaluation is not. A passage can be well drafted and wrong on the law, correct and unsupported by its cited authority, accurate and procedurally inappropriate for the matter. A single quality score collapses those into a number nobody can act on.

02 / BUILDING THE STANDARD

Decomposition into checkable questions.

A workable standard is a bank of specific questions, each with a defined answer space and each narrow enough that two practitioners agree on the answer. MKC2's framework applies 45 such checks to every item. The decomposition is the expensive part and it is also the durable part: it survives model changes, which a prompt does not.

Three properties make a check usable:

  • It is answerable from the document. A check requiring knowledge of the client relationship cannot be scored consistently against a corpus.
  • It occurs often enough to measure. A defect appearing in three items out of four thousand cannot be measured reliably, and should be reported with that caveat rather than averaged in.
  • Practitioners agree on it. Measure inter-rater agreement before treating the labels as ground truth; where experts disagree, the check needs rewriting rather than more data.

03 / THE WORKED EXAMPLE

MKC2, reported in full.

Four measures,
three models,
one run.

Privilege AI's calibration record for MKC2, dated 21 September 2026, measures 45 checks on 4,403 held-out items, with both the model being replaced and the stock Qwen3-8B base re-scored in the same run on the same material.

The MKC2 calibration record. Re-scoring both comparison models in the same run is what makes the differences attributable to the models rather than to the reports.
MeasureMKC2Stock baseModel replaced
Ranking quality0.96230.72810.6457
Answers correct93.17%83.09%80.73%
Brier score0.06200.12740.1409
Calibration error0.06100.13630.1785

Calibration record, 21 September 2026. Benchmark results on held-out evaluation data, not live-matter outcomes. No attorney dispositions have been recorded, so none of these figures show that MKC2 agrees with a lawyer on live work. One check of 45 is refused outright for insufficient ranking quality; 25 carry report-only authority. MKC2 reports; it does not act. Full record on the MKC2 page.

The ranking figure is the one that decides whether the system is usable. A score that orders defective work above sound work only loosely cannot support a review queue, whatever its accuracy: no threshold separates the classes. The model being replaced scored 0.6457 there, below the stock base model, which is the clearest statement of why the work was done.

04 / WHAT A BENCHMARK CANNOT TELL YOU

The limits a practice should insist on hearing.
↳ 01

Whether it agrees with your practitioners.

Held-out evaluation measures behaviour on evaluation data. Agreement on live matters requires a recorded disposition history, which is a separate exercise.

↳ 02

Whether it works on your documents.

Performance is corpus-dependent. Vocabulary, document types and the definition of a defect vary between practices.

↳ 03

Whether the gain survives review capacity.

A system that finds defects nobody has time to read has not improved anything. The review workflow is part of the measurement.

Two questions are worth asking any supplier. Were the baselines re-scored in the same run, or quoted from elsewhere? And which checks failed to reach the bar? A report with no negative findings has either an unusually good model or an unusually incurious author. See AI model benchmarking and held-out evaluation.

05 / QUESTIONS

Asked when assessing a legal AI system.

What accuracy should a legal AI system reach?

+

Accuracy alone is the wrong question. Ask for ranking quality — whether the system orders defective work above sound work — and for calibration, because those decide whether the output can be used to direct review. A high accuracy figure with weak ranking quality cannot support a queue.

How many documents does an evaluation set need?

+

Enough for every reported check to be measurable, which for a decomposed standard means thousands rather than hundreds. MKC2's record uses 4,403 items precisely so that individual checks can be reported separately rather than averaged.

Who should define what counts as a defect?

+

The practitioners who carry the consequence, not the engineering team. A standard written for measurement convenience drifts towards what is easy to score and loses the confidence of the people expected to rely on it.

Can we reuse a vendor's benchmark?

+

As a filter, not as a decision. Performance is corpus-dependent, and a figure produced on someone else's documents with someone else's defect definition says little about your material. Run your own held-out set.

What does report-only authority mean?

+

That a check records a finding for a human reader and changes nothing on its own. It is the tier assigned where a check's measured evidence supports reporting but not gating. Twenty-five of MKC2's 45 checks sit there by design.

Measure it
on your own work.

Privilege AI builds held-out evaluation sets and per-check reporting, and publishes the checks that do not reach the bar.