docketrouter

Rankings

How the models we track score on real, blinded Texas appellate cases, how much of our own traffic each one carries, and what the catalog costs. Every chart here is built from graded runs and logged requests; nothing is projected or illustrative.

Real, blinded Texas appellate cases decided after every model's training cutoff. Score is mechanical: disposition, standard of review, controlling authority and verbatim quotations checked against the real opinion; a fabricated citation or quotation caps the case at 25%. Methodology →

The “with DocketRouter” column is provisional. The “with DocketRouter” scores and every delta on this page are PROVISIONAL and may be overstated. HLL items are built from real Texas appellate opinions, and those same opinions are in the index the grounded arm retrieves from, with nothing excluding them, so a model could read the court's own decision before deciding the case. We found this on 2026-08-28 and halted all HLL testing the same day rather than keep publishing against it. The raw column is unaffected: a raw arm does no retrieval, so it cannot leak. The product is unaffected too, and for a concrete reason rather than reassurance: a live dispute has no written opinion yet, so a customer's query has nothing to leak from. Before these numbers are presented as final, source opinions must be excluded from the grounded arm and the affected items re-run.

A score here is the model and the route it drew. A model id is a routing pool, not one endpoint: the same id is served by many upstream providers and which one answers varies call to call. Measured on our own results, one model's mean grade ranges from 0.22 to 0.54 depending purely on which upstream served it, because the blank-answer rate ranges from 27% to 78% and a blank scores zero against the model. Removing every blank does not remove the effect: an 11.7-point spread survives among non-blank answers, so a route can serve a worse answer without ever returning an empty one. Scores here do not separate the model from the draw. Single-provider ids are unaffected, and you cannot tell which case a model is in unless the route was recorded.

#ModelProviderRawWith DocketRouterΔRank moveCases graded (raw / with DocketRouter)Fabricated citations (raw / with DocketRouter)
1Gemini 3.7 Flash
google/gemini-3.7-flash
Google58.6%64.2%+5.6NEW384 / 3840 / 0
2Solar Pro 4
upstage/solar-pro4
upstage39.2%50.7%+11.5NEW909 / 90923 / 0
3DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash
DeepSeek42.5%50.2%+7.7▼ 2999 / 9996 / 0
4gpt-oss-120b
openai/gpt-oss-120b
OpenAI24.5%38.0%+13.5NEW960 / 9600 / 0
5gpt-oss-20b
openai/gpt-oss-20b
OpenAI2.2%25.8%+23.6NEW659 / 6590 / 0

Latency

Per-item wall time, raw. Dot is the median, ring is the 90th percentile, line spans the full min-max range.

0ms500000ms1000000ms1500000msGemini 3.7 FlashGemini 3.7 Flash median: 19298ms (n=1000)Gemini 3.7 Flash p90: 26133ms (n=1000)Gemini 3.7 Flash: min 8887ms, median 19298ms, p90 26133ms, max 46360ms (n=1000)DeepSeek V4 Flash 0423DeepSeek V4 Flash 0423 median: 43939ms (n=2213)DeepSeek V4 Flash 0423 p90: 180002ms (n=2213)DeepSeek V4 Flash 0423: min 10262ms, median 43939ms, p90 180002ms, max 1241734ms (n=2213)gpt-oss-20bgpt-oss-20b median: 50885ms (n=1000)gpt-oss-20b p90: 138389ms (n=1000)gpt-oss-20b: min 3065ms, median 50885ms, p90 138389ms, max 180040ms (n=1000)gpt-oss-120bgpt-oss-120b median: 55686ms (n=1000)gpt-oss-120b p90: 103353ms (n=1000)gpt-oss-120b: min 1404ms, median 55686ms, p90 103353ms, max 180011ms (n=1000)Solar Pro 4Solar Pro 4 median: 61561ms (n=1000)Solar Pro 4 p90: 142385ms (n=1000)Solar Pro 4: min 3662ms, median 61561ms, p90 142385ms, max 185886ms (n=1000)
median p90 min-max range

Price vs. score

Mean measured price per graded call (raw), including the DocketRouter margin, against mean score. Measured from the actual upstream charge, not the list price.

0%20%40%60%80%$0.0000$0.0020$0.0040$0.0060$0.0080$0.01price per callscoreGemini 3.7 Flash: $0.0099, 62.1% · 386 priced calls
Gemini 3.7 Flash