Defects caught per hour of review.
The metric that reflects the actual claim. Time saved without a defect measure says nothing about whether quality held.
MKC2 results ↗AI is strong at finding and structuring what a document says. Deciding whether what it says is acceptable is a different task with a different error profile.
01 / DEFINITION
The short answer first.AI legal document review is the use of models to read legal documents and produce structured output: extracted terms, identified clauses, flagged deviations or drafted comments. Its usefulness depends on whether the task is extraction, which models do well, or judgement, which is harder to verify.
The distinction matters because the two have different error profiles. An extraction error is usually checkable — the clause either says that or it does not. A judgement error is a view about whether something is acceptable, and verifying it requires the same expertise that would have been needed to form it.
02 / TASK BY TASK
Where AI currently earns its place.| Task | Type | Realistic assessment |
|---|---|---|
| Extracting dates, parties and defined terms | Extraction | Reliable, and verifiable against the source |
| Locating clauses by concept rather than wording | Extraction | Strong; a clear improvement on keyword search |
| Comparing a clause against a standard position | Judgement | Useful for triage; needs review of the conclusion |
| Identifying missing provisions | Judgement | Absence is harder than presence; expect misses |
| Assessing commercial or litigation risk | Judgement | Requires context the document does not contain |
| Verifying that citations resolve | Extraction | Mechanical and high value where an authoritative index exists |
Citation verification deserves more attention than it gets. Where the practice holds an authoritative index — of matters, authorities, contracts or precedent — checking that every cited identifier resolves is close to free and eliminates the most damaging class of fabrication outright. It should be the first thing built, not the last. See AI hallucination detection.
Missing-provision detection is the task most often over-sold. Detecting the absence of something requires knowing what should have been present, which depends on the deal, the jurisdiction and the client's position — none of which is in the document. Treat it as a prompt for attention rather than a finding.
03 / PRIORITISING RATHER THAN REPLACING
The shape of benefit that holds up.The reliable gain is not reviewing less. It is reviewing in a better order, so a fixed amount of expert attention catches more.
This is the argument for supervision rather than automation, and it depends on one property: whether the system's score orders defective work above sound work. MKC2 achieves 0.9623 on that measure on 4,403 held-out items, against 0.7281 for the stock base model it was fine-tuned from — which is what makes an ordered review queue meaningful rather than arbitrary.
Two design details determine whether the ordering translates into caught defects. The confidence attached to each flag has to be calibrated, or reviewers learn to discount it and the ordering information is lost. And the queue has to be capped at reviewable volume: a list of four hundred flags is not a prioritisation, it is a backlog. See human-in-the-loop AI and AI model calibration.
04 / MEASURING THE BENEFIT
What to record before and after.The metric that reflects the actual claim. Time saved without a defect measure says nothing about whether quality held.
Accepted, rejected or unclear, with the reviewer recorded. This is the only data that measures the system on live work.
A tracked share of items with deliberately introduced problems, which measures whether review attention is real.
Privilege AI states plainly that no attorney dispositions have been recorded for MKC2, so its published figures describe held-out evaluation rather than agreement on live matters. Building that disposition record is how a practice turns a benchmark into evidence about its own work — and it is the same record that would eventually justify any increase in the system's authority. See legal AI evaluation.
05 / QUESTIONS
Asked by teams introducing AI into review.For extraction tasks it can do most of the mechanical work, with verification against the source. For judgement tasks the defensible pattern is prioritisation: the system orders the work and a practitioner reviews, which improves the yield of a fixed amount of review rather than removing it.
Good and corpus-dependent, and the useful property is that it is verifiable: an extracted clause can be checked against the source in seconds. That checkability is what makes extraction a safer first deployment than judgement.
It catches different things. Models are consistent across long documents where human attention degrades, and they miss context a practitioner has from the matter. The gain comes from the combination, measured as defects caught per hour of review.
A fluent, well-structured output that is wrong, accepted because it looks like competent work. Calibrated confidence, verified citations and a review queue capped at reviewable volume are the practical defences.
Defects caught per hour of review, before and after, with dispositions recorded on every flag and a share of seeded known defects to confirm that review attention is real. Time saved alone does not establish that quality held.
Privilege AI develops supervision models that rank defective work above sound work, with authority set by measured evidence.