LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
Limitations Arithmetic. The limitations period and the trigger date are stated in the prompt; the model must do the date math (including leap years and discovery-rule triggers) and say whether filing was timely. Isolates reasoning from jurisdiction-specific memorization. Items →
| # | Model | Provider | Raw | Juiced | Δ | Latency | Task cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4.3 x-ai/grok-4.3 | xAI | 100% | 100% | +0 | 4951ms | $0.0362 | $1.25 |
| 2 | Gemini 2.5 Pro google/gemini-2.5-pro | 100% | 100% | +0 | 5977ms | $0.0646 | $1.25 | |
| 3 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 63% | 88% | +25 | 1324ms | $0.0106 | $0.55 |
| 4 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 50% | 88% | +38 | 4744ms | $0.0810 | $3 |
| 5 | Qwen3 32B qwen/qwen3-32b | Qwen | 38% | 75% | +38 | 13960ms | $0.0025 | $0.08 |
| 6 | GPT-5 openai/gpt-5 | OpenAI | 25% | 50% | +25 | 8956ms | $0.0533 | $1.25 |