Humanity's Last Lawsuit (HLL) v0.1: Methodology
Goal. Measure whether a model can reason through a real appellate case it has not seen with the parties removed: identify the dispositive issue, state the controlling standard, cite real authority, apply it, and predict the disposition, and not invent cases.
1. The cases ("blind opinions")
- Source: published Texas appellate opinions with full text, drawn from the Supreme Court of Texas, the Court of Criminal Appeals and the courts of appeals, plus rule-based sections authored from verbatim statutory text. Live counts are on the HLL page; this document does not restate them so the two can never disagree.
- Each case yields several questions that probe different decisions the court made: the dispositive issue and disposition, the standard of review, the procedural vehicle or jurisdiction, the remedy, and sub-issues.
- Blinding: parties, judges, product and brand names and identifying dates are removed. Courts, statutes, rules and other cited cases stay.
- The court is the answer key. Every gold field (disposition, standard of review, the authorities the court relied on) is taken from the opinion itself and anchored to the court's own words. Only cases whose answer key is fully supported by the opinion text are used.
2. Grading
| Axis | Weight | Method |
|---|---|---|
| outcome | 2 | mechanical: the disposition the court reached against the disposition the answer calls, with partial credit for the correct direction |
| standard | 2 | mechanical: the standard of review the court applied against the one the answer names. Graded only where the opinion itself states one; otherwise excluded from that item's denominator |
| authority | 2 | mechanical: every citable authority in the answer (reporter cites, rules, statutes, constitutions) matched against the authorities the court relied on, with the controlling authority weighted most |
| quote fidelity | 2 | mechanical: every quotation attributed to a source must appear verbatim in that source. A quotation attributed to the blinded opinion that is not in it is fabricated and caps the item like a fake citation; a quotation against material we do not hold in full is unverifiable and carries no penalty. Answers with no verifiable quotation have this axis excluded |
| hallucination | gate | every citation in the answer must appear in the opinion or in our citation library; any citation affirmatively shown not to exist caps the item score at 0.25 |
Score = weighted mean in [0,1], then the hallucination cap. No model grades another model: every axis is a deterministic function of (item, answer, library). The verifier is asymmetric: a citation counts as fake only when there is affirmative evidence it does not exist; a citation that merely cannot be verified is never penalized, because penalizing a real case is worse than missing a fake one.
2a. Question formats
Every case yields several asks in different formats: free-form appellate analysis, multiple choice on the disposition, a strict JSON object, the controlling citation alone, a check on whether a quotation is verbatim, and a filing-deadline computation under a named Texas rule. Each format is graded mechanically against a fixed answer key with no model in the loop.
2b. Determinism
The headline score is fully deterministic: outcome, standard, authority, quote fidelity and the hallucination gate are mechanical functions of the item, the answer and the citation library. Same answer, same score, forever, and anyone can recompute it from the published grades. The grader version is stamped on every result.
3. Calibration
Before any leaderboard is published the grader is calibrated on the bank: the court's own answer must score near full marks and a deliberately wrong answer must score near zero.
4. Modes
Every model is run once alone (system prompt only, no retrieval) and once with DocketRouter (retrieval and citation verification injected). Every model gets the same items under the same conditions. The delta between the two arms is the product claim.
5. Limitations (v0.1)
- All items are Texas appellate opinions; see the live counts on the HLL page rather than a number quoted here. All source opinions are from 2026 (post-training-cutoff for most models, which helps against contamination but limits breadth for now).
- Fluent wrong answers exist; the deterministic axes (outcome, authority, hallucination) anchor the score.
- No citation library is complete for every reporter; hence unverified citations are never penalized.