Show your work · Transparency
How MikeCheck measures itself
The goal is 100% agreement with a professional citator on the question that matters — is this case still good law? We are not there. This page is the documented starting line of the climb to get there: what MikeCheck checks, where its honest baseline sits today, and the instruments that will turn each proxy number into a real, published one.
Last reviewed 2026-07-07 · assisted research, not legal advice.
What we measure, and how
MikeCheck's core is a treatment system— the machinery that answers “is this case still good law?” It has two layers: a deterministic engine that reads citing opinions and classifies how they treat the target case, and an evidence-gated, human-reviewed curation pipeline — a registry of known overrulings and abrogations that pins the verdict for landmark cases the engine cannot yet resolve on its own.
A verdict is therefore either registry-assisted (a curated row pins it) or registry-blind(the raw engine's own reading). These two paths perform very differently, and conflating them is the most common way legal-tech tools oversell. We keep them separate everywhere.
The metric stack
Five metrics, behind two north stars — raw parity (match a citator on any case) and calibrated honesty (never be confidently wrong; abstain and escalate instead). The honest baseline as of 2026-07-06:
| # | What it measures | Honest baseline | Status |
|---|---|---|---|
| M1 | Verdict miss rate — an overruled case wrongly called good law (registry-blind) | raw-engine recall 26.0% on the LoC census (extraction-level — see note); frontier-LLM ceiling ~92.8% | proxy — real registry-blind verdict baseline pending arm-2 |
| M2 | Verdict false-alarm rate — a good case wrongly called not-good | 0 / 25 on a landmark set; thin-citation tail unmeasured | partial |
| M3 | Treatment-list precision — is the list of signals an agent reads accurate? | ~15% precision on good-law landmarks (high false-positive rate) | measured, low — a known gap |
| M4 | Calibration — does stated confidence match empirical accuracy? Does it abstain honestly? | unmeasured as a curve | not yet built |
| M5 | Agreement with Westlaw KeyCite (registry-blind sample) | 94.4% exists but is registry-assisted (n=54 = the curated benchmark itself); registry-blind unmeasured | proxy |
The calibrated-honesty posture
A confidently wrongverdict is worse for a user than an honest “I'm not sure.” MikeCheck is built to abstain rather than guess:
- Every case-bearing verdict carries a confidence score and source provenance — citation, source URL, confidence interval, treatment basis, and disqualifying authorities. Never a black box.
- Absence of a negative signal is not asserted as “good law.” When there is citing data but no treatment either way, the verdict reads “NO NEGATIVE SIGNAL FOUND”; when there is too little data to classify, it reads “INSUFFICIENT DATA.”Only affirmative positive treatment earns “GOOD LAW.”
- Coverage is disclosed— verdicts say “analyzed N of M citing cases” and warn when the sample is small.
Calibrated honesty is a north star with, today, no measuring stick — the M4 reliability curve is not yet built. We say so here rather than imply the calibration is proven.
The open substrate
The auditable substrate behind these numbers lives in the open repository — the LoC census and coverage tooling, the Westlaw-agreement methodology and per-case results, the curated registry and its evidence-gated promotions, the Wilson-interval math used so small-n proportions are reported as intervals, the product thesis and metric definitions, and the claim inventory recording every claim retired, kept, or reworded and the instrument that earns each retired one back. “Show your work” means these are published, not paraphrased.
Where this is going
Each instrument below turns a proxy or missing number into a real, published one. As each lands it becomes a dated milestone, and if it validates a claim, that claim can return to the public surfaces under the one rule above.
- Registry-blind end-to-end verdict benchmark (arm-2) — the flagship. Turns M1/M5 from proxy to real; replaces the extrapolated “~74% miss.”
- Deployable-prompt detection measure — does a shippable prompt reach the ~92.8% ceiling, or only a frontier ensemble?
- Treatment-list precision measure (real M3) — span-grounded precision of the signal list agents read.
- Thin-citation false-alarm slice — the M2 tail the landmark set cannot reach.
- M4 calibration curve — a reliability diagram + abstention-honesty measure; the first measuring stick for calibrated honesty.
- State-segment benchmark — hand-curated state-overruling ground truth (~55% of real queries are state appellate).
- Truth-based quote-verification corpus — independent, adversarial labels with a frozen answer-key digest.
- Live-production verdict telemetry — a daily golden-citation canary spanning good-law and non-registry-detectable negatives.
This is a measured climb from a rigorously honest floor. That is the whole point.