LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
Hearsay Identification. Given a short fact pattern, decide whether the out-of-court statement is being offered for the truth of the matter asserted (hearsay) or for a non-hearsay purpose (effect on listener, verbal act, state of mind, impeachment). Modeled on the LegalBench hearsay task. Items →
| # | Model | Provider | Raw | Juiced | Δ | Latency | Task cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 80% | 100% | +20 | 1422ms | $0.0110 | $0.55 |
| 2 | Grok 4.3 x-ai/grok-4.3 | xAI | 100% | 100% | +0 | 3281ms | $0.0357 | $1.25 |
| 3 | Gemini 2.5 Pro google/gemini-2.5-pro | 100% | 100% | +0 | 6106ms | $0.0748 | $1.25 | |
| 4 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 100% | 100% | +0 | 4487ms | $0.0847 | $3 |
| 5 | GPT-5 openai/gpt-5 | OpenAI | 70% | 90% | +20 | 6637ms | $0.0621 | $1.25 |
| 6 | Qwen3 32B qwen/qwen3-32b | Qwen | 80% | 50% | -30 | 17815ms | $0.0029 | $0.08 |