Privilege AIMKC2 results

A demanding
application.

Legal work combines confidential material, a high cost of plausible error and a verifiable citation standard. It is where the weaknesses of AI systems show first.

READ WHY ↓PRIVILEGE · SUPERVISION · VERIFIABILITY

01 / DEFINITION

The short answer first.

What is legal AI?

Legal AI is the application of AI systems to legal work: drafting, review, research, extraction and analysis. It is distinguished from general AI applications by material that is confidential by default, a professional standard for verifiable citation, and a cost of error borne by a client.

Privilege AI's work in this domain is concrete rather than general. MKC2, a supervision model fine-tuned from Qwen3-8B, was built for HIGHCOURT to evaluate AI-generated legal work and the actions of legal AI agents, and it runs entirely offline. The pages in this cluster describe what that application requires, not what legal AI could theoretically do.

02 / FOUR PROPERTIES THAT CHANGE THE DESIGN

Why general guidance does not transfer.
01 — PRIVILEGE

Confidential
by default.

Client material is not merely sensitive; confidentiality is a professional obligation. That makes the inference boundary a first-order design constraint rather than a preference.

OfflineNo egress
02 — PLAUSIBILITY

Wrong
reads as right.

A defective clause or an unsupported proposition is fluent, well formed and professionally phrased. The failure mode is not visible in the output.

ReviewSupervision
03 — VERIFIABILITY

Citations
must resolve.

Legal argument rests on sources that either exist and say what is claimed, or do not. That makes a large share of errors mechanically checkable.

AttributionSource checks

A fourth property is asymmetric review capacity. AI-assisted drafting increases output faster than the senior attention available to check it, and the gap is where a defect reaches a client. That imbalance is the specific problem a supervision model addresses: not replacing review, but ordering it so that scarce expert attention reaches the work most likely to be wrong.

03 / WHAT PRIVILEGE AI HAS BUILT

First-party evidence.

AI that questions
AI.

MKC2 is a specialised supervision model that reviews AI-generated legal drafting and the actions of legal AI agents. On 4,403 held-out items, with 45 separate checks applied to each, it ranks defective work above sound work at 0.9623, against 0.7281 for the stock base model it was fine-tuned from and 0.6457 for the system it replaces — all three re-scored in the same run on the same material. Its accuracy is 93.17% against 83.09% for the stock base.

The record is equally explicit about what it does not show. These are benchmark results on held-out evaluation data, not live-matter outcomes. No attorney dispositions have been recorded, so nothing in the record demonstrates that MKC2 agrees with a lawyer on live work. MKC2 operates in review-only mode: it reports and flags, and it does not act. One of the 45 checks is refused outright for insufficient ranking quality and 25 carry report-only authority.

The full record, including Brier score and calibration error against both comparison models, is on the MKC2 page.

04 / WHERE TO BE CAUTIOUS

Honest limits on what these systems do.
↳ 01

Benchmarks are not practice.

A held-out evaluation measures behaviour on evaluation data. Agreement with a practitioner on live matters is a separate claim requiring separate evidence.

↳ 02

Professional obligations are not a model property.

Duties of competence, confidentiality and supervision rest with the practice. A system can support them; it cannot discharge them.

↳ 03

Specialisation is narrow by design.

A model fine-tuned for one review task is not general legal capability, and should not be described as if it were.

Rules of professional conduct differ by jurisdiction and are a matter for the practice rather than for a technology supplier. What a supplier can legitimately provide is a system whose data boundary, measured performance and authority limits are stated precisely enough to be assessed — which is why Privilege AI publishes what a check cannot do alongside what it can. See legal AI evaluation and AI model supervision.

05 / QUESTIONS

Asked by practices assessing AI systems.

What makes legal AI harder than other applications?

+

Three properties together: material that is confidential by default, errors that are fluent and professionally phrased rather than obviously wrong, and a citation standard that must resolve to real sources. Each is manageable alone; combined, they raise the requirement on both the data boundary and the review process.

Can legal AI run without sending data to a third party?

+

Yes. Open-weight models fine-tuned for specific legal tasks run on local hardware, and MKC2 runs entirely offline for exactly this reason. See private legal AI.

Does legal AI replace junior review work?

+

The evidence supports a different shape: ordering review so that expert attention reaches the riskiest work first. A supervision model that ranks defective work above sound work improves the yield of a fixed amount of review rather than removing the need for it.

How should a practice evaluate a legal AI system?

+

On its own documents, with a held-out set, reporting ranking quality and calibration alongside accuracy, and with baselines re-scored in the same run. Vendor benchmark figures on public datasets say very little about performance on a particular practice's material. See legal AI evaluation.

What should a legal AI system never do?

+

Take an irreversible action on a matter without a person confirming it. MKC2's review-only ceiling is a deliberate design decision rather than a current limitation: the cost of a wrong autonomous action on a live matter is not symmetrical with the cost of a wrong flag.

Built for
a demanding domain.

Privilege AI develops specialised models and private systems for work where confidentiality and defensible measurement both matter.