← Back to home

Show your work · Transparency

How MikeCheck measures itself

The goal is 100% agreement with a professional citator on the question that matters — is this case still good law? We are not there. This page is the documented starting line of the climb to get there: what MikeCheck checks, where its honest baseline sits today, and the instruments that will turn each proxy number into a real, published one.

Last reviewed 2026-07-07 · assisted research, not legal advice.

The one rule
No performance number goes on a public surface ahead of the instrument that validates it. When a claim returns, it returns tied to a named measurement, disclosed in both registry-assisted and registry-blind form where that distinction applies, and logged as a dated entry in the progress log. Until then MikeCheck describes what it does — not how well it does it.

What we measure, and how

MikeCheck's core is a treatment system— the machinery that answers “is this case still good law?” It has two layers: a deterministic engine that reads citing opinions and classifies how they treat the target case, and an evidence-gated, human-reviewed curation pipeline — a registry of known overrulings and abrogations that pins the verdict for landmark cases the engine cannot yet resolve on its own.

A verdict is therefore either registry-assisted (a curated row pins it) or registry-blind(the raw engine's own reading). These two paths perform very differently, and conflating them is the most common way legal-tech tools oversell. We keep them separate everywhere.

The metric stack

Five metrics, behind two north stars — raw parity (match a citator on any case) and calibrated honesty (never be confidently wrong; abstain and escalate instead). The honest baseline as of 2026-07-06:

#What it measuresHonest baselineStatus
M1Verdict miss rate — an overruled case wrongly called good law (registry-blind)raw-engine recall 26.0% on the LoC census (extraction-level — see note); frontier-LLM ceiling ~92.8%proxy — real registry-blind verdict baseline pending arm-2
M2Verdict false-alarm rate — a good case wrongly called not-good0 / 25 on a landmark set; thin-citation tail unmeasuredpartial
M3Treatment-list precision — is the list of signals an agent reads accurate?~15% precision on good-law landmarks (high false-positive rate)measured, low — a known gap
M4Calibration — does stated confidence match empirical accuracy? Does it abstain honestly?unmeasured as a curvenot yet built
M5Agreement with Westlaw KeyCite (registry-blind sample)94.4% exists but is registry-assisted (n=54 = the curated benchmark itself); registry-blind unmeasuredproxy
Read the 26% precisely
The 26% is an extraction-recallnumber, not a verdict-accuracy number. It is the raw engine's recall on arm 1 of the Library-of-Congress overruling census — how often the deterministic extractor surfaces a known overruling signal from citing text, measured with no registry assistance. It is not the end-to-end verdict accuracy a user experiences (production verdicts are registry-assisted for landmark cases), and it is notthe registry-blind verdict baseline. That head-to-head number — census arm 2 — is not yet built. Verdict baseline: pending arm-2.We switched off the registry and every other assist to measure the raw extraction layer alone, so the number is that layer's true floor rather than one flattered by the parts that already work. That floor is the honest starting line we measure up from, on purpose.
Read the 94.4% precisely
The 94.4% Westlaw-KeyCite agreement figure is real and reproducible, but it is registry-assisted, n=54, where the 54 cases are the curated benchmark itself. It grades the curated registries against KeyCite on hand-selected landmark cases — not raw-engine accuracy and not coverage across all law. With n=54 the Wilson interval around any such proportion is wide: it is a spot check, not a population accuracy claim.

The calibrated-honesty posture

A confidently wrongverdict is worse for a user than an honest “I'm not sure.” MikeCheck is built to abstain rather than guess:

  • Every case-bearing verdict carries a confidence score and source provenance — citation, source URL, confidence interval, treatment basis, and disqualifying authorities. Never a black box.
  • Absence of a negative signal is not asserted as “good law.” When there is citing data but no treatment either way, the verdict reads “NO NEGATIVE SIGNAL FOUND”; when there is too little data to classify, it reads “INSUFFICIENT DATA.”Only affirmative positive treatment earns “GOOD LAW.”
  • Coverage is disclosed— verdicts say “analyzed N of M citing cases” and warn when the sample is small.

Calibrated honesty is a north star with, today, no measuring stick — the M4 reliability curve is not yet built. We say so here rather than imply the calibration is proven.

The open substrate

The auditable substrate behind these numbers lives in the open repository — the LoC census and coverage tooling, the Westlaw-agreement methodology and per-case results, the curated registry and its evidence-gated promotions, the Wilson-interval math used so small-n proportions are reported as intervals, the product thesis and metric definitions, and the claim inventory recording every claim retired, kept, or reworded and the instrument that earns each retired one back. “Show your work” means these are published, not paraphrased.

Where this is going

Each instrument below turns a proxy or missing number into a real, published one. As each lands it becomes a dated milestone, and if it validates a claim, that claim can return to the public surfaces under the one rule above.

  1. Registry-blind end-to-end verdict benchmark (arm-2) — the flagship. Turns M1/M5 from proxy to real; replaces the extrapolated “~74% miss.”
  2. Deployable-prompt detection measure — does a shippable prompt reach the ~92.8% ceiling, or only a frontier ensemble?
  3. Treatment-list precision measure (real M3) — span-grounded precision of the signal list agents read.
  4. Thin-citation false-alarm slice — the M2 tail the landmark set cannot reach.
  5. M4 calibration curve — a reliability diagram + abstention-honesty measure; the first measuring stick for calibrated honesty.
  6. State-segment benchmark — hand-curated state-overruling ground truth (~55% of real queries are state appellate).
  7. Truth-based quote-verification corpus — independent, adversarial labels with a frozen answer-key digest.
  8. Live-production verdict telemetry — a daily golden-citation canary spanning good-law and non-registry-detectable negatives.

This is a measured climb from a rigorously honest floor. That is the whole point.