docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

Limitations Arithmetic. The limitations period and the trigger date are stated in the prompt; the model must do the date math (including leap years and discovery-rule triggers) and say whether filing was timely. Isolates reasoning from jurisdiction-specific memorization. Items →

#ModelProviderRawJuicedΔLatencyTask costInput $/M
1Grok 4.3
x-ai/grok-4.3
xAI100%100%+04951ms$0.0362$1.25
2Gemini 2.5 Pro
google/gemini-2.5-pro
Google100%100%+05977ms$0.0646$1.25
3DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek63%88%+251324ms$0.0106$0.55
4Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic50%88%+384744ms$0.0810$3
5Qwen3 32B
qwen/qwen3-32b
Qwen38%75%+3813960ms$0.0025$0.08
6GPT-5
openai/gpt-5
OpenAI25%50%+258956ms$0.0533$1.25