Rankings
How the models we track score on real, blinded Texas appellate cases, how much of our own traffic each one carries, and what the catalog costs. Every chart here is built from graded runs and logged requests; nothing is projected or illustrative.
Real, blinded Texas appellate cases decided after every model's training cutoff. Score is mechanical: disposition, standard of review, controlling authority and verbatim quotations checked against the real opinion; a fabricated citation or quotation caps the case at 25%. Methodology →
The “with DocketRouter” column is provisional. The “with DocketRouter” scores and every delta on this page are PROVISIONAL and may be overstated. HLL items are built from real Texas appellate opinions, and those same opinions are in the index the grounded arm retrieves from, with nothing excluding them, so a model could read the court's own decision before deciding the case. We found this on 2026-08-28 and halted all HLL testing the same day rather than keep publishing against it. The raw column is unaffected: a raw arm does no retrieval, so it cannot leak. The product is unaffected too, and for a concrete reason rather than reassurance: a live dispute has no written opinion yet, so a customer's query has nothing to leak from. Before these numbers are presented as final, source opinions must be excluded from the grounded arm and the affected items re-run.
A score here is the model and the route it drew. A model id is a routing pool, not one endpoint: the same id is served by many upstream providers and which one answers varies call to call. Measured on our own results, one model's mean grade ranges from 0.22 to 0.54 depending purely on which upstream served it, because the blank-answer rate ranges from 27% to 78% and a blank scores zero against the model. Removing every blank does not remove the effect: an 11.7-point spread survives among non-blank answers, so a route can serve a worse answer without ever returning an empty one. Scores here do not separate the model from the draw. Single-provider ids are unaffected, and you cannot tell which case a model is in unless the route was recorded.
| # | Model | Provider | Raw | With DocketRouter | Δ | Rank move | Cases graded (raw / with DocketRouter) | Fabricated citations (raw / with DocketRouter) |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash google/gemini-3.7-flash | 58.6% | 64.2% | +5.6 | NEW | 384 / 384 | 0 / 0 | |
| 2 | Solar Pro 4 upstage/solar-pro4 | upstage | 39.2% | 50.7% | +11.5 | NEW | 909 / 909 | 23 / 0 |
| 3 | DeepSeek V4 Flash 0423 deepseek/deepseek-v4-flash | DeepSeek | 42.5% | 50.2% | +7.7 | ▼ 2 | 999 / 999 | 6 / 0 |
| 4 | gpt-oss-120b openai/gpt-oss-120b | OpenAI | 24.5% | 38.0% | +13.5 | NEW | 960 / 960 | 0 / 0 |
| 5 | gpt-oss-20b openai/gpt-oss-20b | OpenAI | 2.2% | 25.8% | +23.6 | NEW | 659 / 659 | 0 / 0 |
Latency
Per-item wall time, raw. Dot is the median, ring is the 90th percentile, line spans the full min-max range.
Price vs. score
Mean measured price per graded call (raw), including the DocketRouter margin, against mean score. Measured from the actual upstream charge, not the list price.