LLM Rankings
Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.
62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.
OverallHearsay IdentificationBluebook Citation FormatFederal Civil ProcedureLimitations ArithmeticContract Clause ClassificationCitation Hallucination Resistance
Contract Clause Classification. Classify a contract clause into one of eight categories drawn from the CUAD taxonomy: Governing Law, Non-Compete, Indemnification, Limitation of Liability, Confidentiality, Termination, Assignment, Force Majeure. The core skill behind contract-review products. Items →
| # | Model | Provider | Raw | Juiced | Δ | Latency | Task cost | Input $/M |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3 32B qwen/qwen3-32b | Qwen | 100% | 100% | +0 | 8997ms | $0.0031 | $0.08 |
| 2 | DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | DeepSeek | 100% | 100% | +0 | 3941ms | $0.0144 | $0.55 |
| 3 | Grok 4.3 x-ai/grok-4.3 | xAI | 100% | 100% | +0 | 2593ms | $0.0421 | $1.25 |
| 4 | GPT-5 openai/gpt-5 | OpenAI | 100% | 100% | +0 | 3173ms | $0.0490 | $1.25 |
| 5 | Gemini 2.5 Pro google/gemini-2.5-pro | 100% | 100% | +0 | 6025ms | $0.0931 | $1.25 | |
| 6 | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 100% | 100% | +0 | 3842ms | $0.1021 | $3 |