LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
Federal Civil Procedure. Short-answer questions on the Federal Rules of Civil and Appellate Procedure: which rule governs a motion, how many days a deadline is, numeric limits on discovery. These are the kinds of facts a paralegal must never get wrong. Items →
| # | Model | Provider | Raw | Juiced | Δ | Latency | Task cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 100% | 100% | +0 | 1453ms | $0.0153 | $0.55 |
| 2 | Grok 4.3 x-ai/grok-4.3 | xAI | 100% | 100% | +0 | 2247ms | $0.0432 | $1.25 |
| 3 | GPT-5 openai/gpt-5 | OpenAI | 100% | 100% | +0 | 2911ms | $0.0516 | $1.25 |
| 4 | Gemini 2.5 Pro google/gemini-2.5-pro | 100% | 100% | +0 | 5245ms | $0.0833 | $1.25 | |
| 5 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 100% | 100% | +0 | 2981ms | $0.0932 | $3 |
| 6 | Qwen3 32B qwen/qwen3-32b | Qwen | 83% | 92% | +8 | 10529ms | $0.0032 | $0.08 |