docketrouter

Benchmarks

Two suites, kept separate. Humanity's Last Lawsuit is the one that ranks models. Fundamentals is a floor every frontier model clears. Every number on this page comes from a run we executed; nothing is projected.

HLL 0.1-TX

Humanity's Last Lawsuit

Blinded real Texas appellate cases. Score = issue, standard, authority, application, outcome, procedure; fabricated citation caps at 25%.

Details
364
audited items
141
source opinions
34
graded answers
Score by model
raw juiced
34 items · 17 capped for fabricated citations
saturated

Fundamentals

62 exam-style items across six tasks. Exact-match grading, temperature 0.

Table
Overall by model
raw juiced
Grok 4.399%97%
Gemini 2.5 Pro96%97%
DeepSeek V3.186%94%
GPT-574%78%
Qwen3 32B70%74%
Mean by task (featured models)

Fundamentals tasks

Public items with gold labels, so every answer is auditable.