docketrouter

Testing your model on HLL

Humanity's Last Lawsuit (HLL) is a bank of real Texas appellate cases with the parties blinded. A model reads the facts and the question, and we check its answer against what the court actually did. This page explains how outside models get tested and what we keep secret.

Two piles of cases

Public pile (20% of the bank). The cases, the questions, and the answer key are published. Anyone can download them and run their own model for free. Use this to test yourself and to check our grader.

Private pile (80% of the bank). The cases and the answer key never leave our server. Only DocketRouter runs models against this pile. This is what the leaderboard is scored on, so nobody can train on the test.

Every new version of HLL rotates cases between the two piles, and case ids carry no information about which pile a case came from.

What we keep secret

  • The private cases and their answer keys.
  • Nothing else. The grading method and the public pile are open, and every published score can be recomputed from the published grades.

How grading works

No model grades another model. Every score is computed mechanically from the answer, the real opinion, and our citation library. Every model gets the same cases under the same conditions:

  • Outcome. Did the answer call the disposition the court reached (affirm, reverse, remand, dismiss...).
  • Standard of review. Where the court stated one (de novo, abuse of discretion, sufficiency...), did the answer name the same one.
  • Authority. Did the answer cite the authority the court relied on.
  • Quote fidelity. Every quotation in the answer must exist word for word in the source it is attributed to.
  • Fabrication gate. Every citation must exist. A made-up citation or a made-up quotation caps the case at 25%.

Same answer, same score, every time. Anyone can recompute it.

How a submission runs

  1. Sign in, give your model a name, an OpenAI-compatible endpoint (or an OpenRouter model id), and a version string.
  2. We call your endpoint from our server for every private case, store the answers, and grade them.
  3. You get an overall score, scores by category, the number of cases capped for fabrication, and a pass/fail per case (never the case text). The result appears on the leaderboard as community · <name> with the date and version.
  4. One official run per model version every 30 days. Your endpoint must answer at least 95% of the cases or the run is void.

Keeping it honest

  • Private cases have never been seen by any public model, and they rotate every version.
  • Cases are sent in random order, so replayed answers stand out.
  • We keep every raw answer. Disputes get a human re-check of the disputed cases.