LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
Bluebook Citation Format. Four variants of a citation to a well-known case or statute; only one follows Bluebook form (reporter abbreviation, volume/page order, parenthetical year, section symbol). Tests fine-grained formatting discipline that matters in filed briefs. Items →
| # | Model | Provider | Raw | Juiced | Δ | Latency | Task cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4.3 x-ai/grok-4.3 | xAI | 100% | 100% | +0 | 4032ms | $0.0368 | $1.25 |
| 2 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 0% | 100% | +100 | 7461ms | $0.1130 | $3 |
| 3 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 100% | 88% | -12 | 1542ms | $0.0117 | $0.55 |
| 4 | Gemini 2.5 Pro google/gemini-2.5-pro | 100% | 88% | -12 | 7538ms | $0.0686 | $1.25 | |
| 5 | GPT-5 openai/gpt-5 | OpenAI | 100% | 75% | -25 | 7796ms | $0.0579 | $1.25 |
| 6 | Qwen3 32B qwen/qwen3-32b | Qwen | 75% | 63% | -12 | 12114ms | $0.0029 | $0.08 |