Privilege AIResearch notes

PRIVILEGE AI / RESEARCH & ENGINEERING

Intelligence,
under scrutiny.

We develop specialised models and study how they reason,
retain knowledge, and evaluate the work of other models.

SCROLL TO EXPLORE ↓MODELS · KNOWLEDGE · SYSTEMS

01 / MODEL RESEARCH

Developed, fine-tuned and evaluated in-house.
MKCSUPERVISION MODEL / BUILT FOR HIGHCOURT
PRIVILEGE AIMODEL SUPERVISIONPRIVATE INFERENCE

AI that questions
AI.

MKC is a specialised supervision model developed by Privilege AI: a model that evaluates AI-generated work. It runs entirely offline.

MKC is the first fine-tuned model for law to cut false alarms on clean legal drafting from 62% to 5.9% — measured against the model it was fine-tuned from, on the same sections, with no data leaving the building.

Explore the measured results

MEASURED. NOT ASSUMED.

Fewer false alarms. Measured.

INITIAL EVALUATION · 03 SEP 2026
62%5.9%

False-positive rate on clean legal drafting.
116 false alarms reduced to 11, across the same 187 clean sections.

unadapted Qwen3-8B62.0%
MKC5.9%
90.5%

Relative reduction in false positives

Based on the recorded counts:
116 versus 11 false alarms.

696

Paired evaluation items

Both models evaluated in the same run.
Initial internal benchmark.

0 / 90

Wrongful tool-block verdicts

In the initial evaluation.
Unsafe-call recall: 56.7%.

Initial internal evaluation, 3 September 2026, against unadapted Qwen3-8B, the base model MKC is fine-tuned from. Legal defect recall was 77%. MKC is deployed and running in HIGHCOURT; the figures here are from controlled evaluation on 187 clean sections, not from live matters. MKC reports; it does not act.

Results in context +

In the initial internal evaluation on 3 September 2026, unadapted Qwen3-8B flagged 116 of 187 clean legal sections; MKC flagged 11. The false-positive rate fell from 62.0% to 5.9%, a 90.5% relative reduction.

Legal defect recall was 77% for MKC and 89% for the baseline. Fewer false alarms came with lower defect recall. These figures describe this benchmark, not general accuracy or live-matter performance.

MKC runs in review-only mode: it reports and flags, it does not block or act on its own. That boundary is deliberate. The benchmark measures review quality, not authority to act.

02 / RESEARCH DIRECTIONS

Questions we are working on.

MODEL SUPERVISION

When should a model
trust another model?

We study the reliability of AI-generated work and the role of specialised models in reviewing it. Our work in HIGHCOURT brings these questions into a demanding application.

Read the evaluation note ↗

EVALUATION & AUTHORITY

Capability is measured.
Authority is earned.

Strong benchmark performance is one part of responsible deployment. We distinguish measured capability from the responsibilities a model should be given.

The reported model operates in review-only mode.

03 / RESEARCH & ENGINEERING

The work spans the full system.

Models.
Knowledge.
Infrastructure.

We work across model adaptation, retrieval and private deployment. Each layer changes what a system can do, what it can know, and how its behaviour can be evaluated.

01 — MODELS

Training and
fine-tuning.

Dataset construction, model training and fine-tuning, with task-specific evaluation and held-out comparisons.

Model trainingFine-tuningEvaluation
02 — KNOWLEDGE

Retrieval and
knowledge.

Retrieval-augmented generation and knowledge systems. Research into the relationship between context and model performance.

RAGData pipelinesSearch
03 — SYSTEMS

Inference and
systems.

Local inference, application integration and on-premise infrastructure. Security boundaries considered throughout the system.

AI engineeringIntegrationOn-prem

04 / PRIVATE SYSTEMS

Control is part of the architecture.

Private inference.
Explicit boundaries.

Deployment is part of the research problem. Data access, model authority and operational constraints shape how AI behaves outside a benchmark.

↳ 01

On-premise inference.

Local model execution with explicit hardware, network and operational constraints.

↳ 02

Access with intention.

Permissions and retrieval boundaries define which information a model can access and which actions it can propose.

↳ 03

Security from the start.

Sensitive data, model behaviour and deployment risk are considered throughout development and evaluation.

CONTROLLED ENVIRONMENT
Data Models Applications
ONE CONTROLLED BOUNDARY

05 / ABOUT PRIVILEGE AI

Experience compounds.

OUR FOUNDATION

Built on decades
of experience.

A continuous practice
of engineering.

Built on decades of experience in software engineering and data science, Privilege AI brings that foundation to the development of specialised models and AI systems.

We connect that experience to the whole system: how data is structured, how software is built, and how intelligence becomes useful.

Web developmentTechnical SEOData scienceAI engineering
Privilege AI

CONTACT

Start a
conversation.

For research, technical enquiries and collaborations.
Tell us a little about what you’re working on.

Please leave confidential material out of this initial message.