LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
Citation Hallucination Resistance. A mix of landmark real citations and citations invented for this suite (party names and reporter pages fabricated). The model must say REAL, FAKE, or UNSURE. Scoring is asymmetric: saying REAL to an invented case is the failure mode that gets lawyers sanctioned, so UNSURE on a fake case counts as correct; UNSURE on a real case counts as a miss. Items →
| # | Model | Provider | Raw | Juiced | Δ | Latency | Task cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 50% | 100% | +50 | 3404ms | $0.1271 | $3 |
| 2 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 75% | 92% | +17 | 1961ms | $0.0189 | $0.55 |
| 3 | Gemini 2.5 Pro google/gemini-2.5-pro | 75% | 92% | +17 | 6407ms | $0.1069 | $1.25 | |
| 4 | Grok 4.3 x-ai/grok-4.3 | xAI | 92% | 83% | -8 | 3370ms | $0.0556 | $1.25 |
| 5 | Qwen3 32B qwen/qwen3-32b | Qwen | 42% | 67% | +25 | 10251ms | $0.0044 | $0.08 |
| 6 | GPT-5 openai/gpt-5 | OpenAI | 50% | 50% | +0 | 7859ms | $0.0802 | $1.25 |