LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
| # | Model | Provider | Raw | Juiced | Δ | Latency | Suite cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 67% | 98% | +31 | 4487ms | $0.6010 | $3 |
| 2 | Grok 4.3 x-ai/grok-4.3 | xAI | 99% | 97% | -1 | 3412ms | $0.2496 | $1.25 |
| 3 | Gemini 2.5 Pro google/gemini-2.5-pro | 96% | 97% | +1 | 6216ms | $0.4913 | $1.25 | |
| 4 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 86% | 94% | +8 | 1941ms | $0.0820 | $0.55 |
| 5 | GPT-5 openai/gpt-5 | OpenAI | 74% | 78% | +3 | 6222ms | $0.3540 | $1.25 |
| 6 | Qwen3 32B qwen/qwen3-32b | Qwen | 70% | 74% | +5 | 12278ms | $0.0190 | $0.08 |