Answers correct
Against 83% for the stock model,
scored on the same material.
MKC2 results ↗MKC2 is a specialised supervision model developed by Privilege AI. It evaluates AI-generated work and the actions of AI agents, and it runs entirely offline.
01 / THE MODEL
What MKC2 is.MKC2 is a supervision model developed by Privilege AI and fine-tuned from Qwen3-8B. It reviews AI-generated legal drafting and the actions of legal AI agents, ranking defective work above sound work so that scarce expert attention goes where the risk is. It runs entirely offline, in review-only mode.
MKC2 was built for HIGHCOURT, where the review problem is concrete: more AI-assisted drafting is produced than can be read line by line, and the cost of a defect reaching a client is high. It is, as far as Privilege AI is aware, the first fully offline, fine-tuned supervision model built specifically to judge legal drafting and legal-agent actions.
Two design decisions define it. It is specialised: the judgement standard for what counts as defective legal work is long, expert and stable, which is the shape of problem fine-tuning suits. And it is bounded: it reports and flags, and it does not block or act on its own.
02 / THE RECORD
Measured, not assumed.Calibration record, 21 September 2026. Every figure is measured on 4,403 held-out items, with the model being replaced re-scored in the same run on the same material.
Put one flawed passage next to one sound passage. MKC2 picks out the flawed one 96 times out of 100.
The stock model it was built from manages 73. The model it replaces, 65.
Against 83% for the stock model,
scored on the same material.
One check could not be shown to work,
so it is switched off, not shipped.
None of them seen during training.
45 separate checks on each.
Calibration record, 21 September 2026. MKC2 is fine-tuned from Qwen3-8B. Every figure is measured on 4,403 held-out items, with the model it replaces re-scored in the same run on the same material. No attorney dispositions have been recorded, so none of this is evidence that MKC2 agrees with a lawyer on live work. MKC2 reports; it does not act.
MKC2 leads on all four measures against both the model it replaces and the stock base model: ranking quality 0.9623 against 0.6457 and 0.7281; Brier score 0.0620 against 0.1409 and 0.1274; calibration error 0.0610 against 0.1785 and 0.1363; correct 93.17% against 80.73% and 83.09%.
The gain is largest on the measure that decides whether a confidence figure is usable at all: how well the model ranks defective work above sound work. The model it replaces scored 0.6457 there, below the stock base model.
Of 45 questions, one is refused outright for insufficient ranking quality and 25 carry report-only authority, changing nothing. Questions that cannot be shown to work are not relied upon.
MKC2 runs in review-only mode: it reports and flags, it does not block or act on its own. That boundary is deliberate. These are benchmark results on held-out data, not live-matter outcomes.
03 / HOW TO READ IT
Which number carries which claim.| Measure | MKC2 | Stock base (Qwen3-8B) | Model replaced | What it licenses |
|---|---|---|---|---|
| Ranking quality | 0.9623 | 0.7281 | 0.6457 | That a threshold on the score can separate defective from sound work |
| Answers correct | 93.17% | 83.09% | 80.73% | That the model is competent at the checks it is given |
| Brier score | 0.0620 | 0.1274 | 0.1409 | That predictions are both right and appropriately confident |
| Calibration error | 0.0610 | 0.1363 | 0.1785 | That a stated confidence can be shown to a reviewer |
The ranking figure is the one to read first. A score with weak ranking quality cannot support any routing rule, however accurate the model is on average — there is no threshold that separates the classes. The model being replaced scored 0.6457 there, below the stock base model, which is the clearest statement of why the work was done.
Definitions for each measure are set out in LLM evaluation metrics, and the reason calibration is reported separately from accuracy is developed in AI model calibration.
04 / THE BOUNDARY
What MKC2 does not do.MKC2 runs in review-only mode. It reports and flags. It does not block work, change a document or take an action on its own.
No attorney dispositions have been recorded. The record measures behaviour on held-out evaluation data, which is a different and weaker claim than agreement on live matters.
One check of 45 is refused outright for insufficient ranking quality, and 25 carry report-only authority. A check that cannot be shown to work is switched off, not averaged in.
These boundaries are part of the design rather than caveats attached to it. The reasoning is set out in AI model supervision: capability is measured, and authority is granted separately, per check, by the organisation carrying the risk.
05 / DEPLOYMENT
How it runs.MKC2 runs entirely offline. Reviewed material, which in a legal setting is privileged by default, never leaves the environment it is held in.
Fine-tuned from Qwen3-8B rather than built from scratch, which keeps the hardware requirement modest and the inference cost predictable.
Each of the 45 checks carries the authority its own measured evidence supports, from refused through report-only to flagging for review.
Deployment considerations for models of this class — hardware, quantisation, throughput and the network boundary — are covered in on-premise LLM deployment and private AI inference. The legal-specific case is set out in private legal AI.
06 / QUESTIONS
Asked about the model and its record.It reads AI-generated work and proposed AI agent actions, applies 45 separate checks to each item, and reports findings with a confidence score. The output is a prioritised set of flags for a human reviewer, not a decision. It does not block or change anything.
Placed against one sound passage, MKC2 scores a defective passage higher about 96 times in 100. That ordering property is what makes a review queue useful: the items most likely to be defective rise to the top. The stock base model it was fine-tuned from scores 0.7281 on the same items.
That has not been measured. No attorney dispositions have been recorded, so the record shows performance on held-out evaluation data and nothing about agreement on live matters. Establishing that would require a sustained disposition record from real review work.
Yes. It runs entirely offline, which is the reason it can be used on privileged material. No prompt, document or output is sent to an external inference service.
Because their measured evidence supports reporting a finding but not gating on it. Authority is granted per check, by the strength of that check's ranking quality and calibration, rather than by the model's overall score. One further check is refused outright and switched off.
MKC2 was developed for HIGHCOURT. For enquiries about supervision models, evaluation work or related research, the fastest route is a direct conversation with the team through the contact form.
The calibration record above is the whole of what has been measured. For the methodology behind it, or for supervision work of your own, start a conversation.