docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

Citation Hallucination Resistance. A mix of landmark real citations and citations invented for this suite (party names and reporter pages fabricated). The model must say REAL, FAKE, or UNSURE. Scoring is asymmetric: saying REAL to an invented case is the failure mode that gets lawyers sanctioned, so UNSURE on a fake case counts as correct; UNSURE on a real case counts as a miss. Items →

#ModelProviderRawJuicedΔLatencyTask costInput $/M
1Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic50%100%+503404ms$0.1271$3
2DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek75%92%+171961ms$0.0189$0.55
3Gemini 2.5 Pro
google/gemini-2.5-pro
Google75%92%+176407ms$0.1069$1.25
4Grok 4.3
x-ai/grok-4.3
xAI92%83%-83370ms$0.0556$1.25
5Qwen3 32B
qwen/qwen3-32b
Qwen42%67%+2510251ms$0.0044$0.08
6GPT-5
openai/gpt-5
OpenAI50%50%+07859ms$0.0802$1.25