docketrouter
Models / OpenAI

o3

by OpenAI · openai/o3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

reasoningtool-usevisionreleased 2025-04-16
Legal score · raw
-
not yet benchmarked
Context
200K
max output 100K
Input
$2
per 1M tokens
Output
$8
per 1M tokens
Suite cost
-
run the suite to see

Benchmark results

TaskCategoryRawJuicedCorrectLatencyCostRan
Hearsay IdentificationEvidence------
Bluebook Citation FormatResearch & Writing------
Federal Civil ProcedureProcedure------
Limitations ArithmeticProcedure------
Contract Clause ClassificationContracts------
Citation Hallucination ResistanceReliability------

Measured by DocketBuster

These numbers come from DocketBuster's own legal battery, not from DocketRouter's suite. Latest run per metric, with n and a 95% Wilson interval where the source reports one. See docketbuster.com/benchmarks.

MetricValuenIntervalMeasured
Statute pinpoint, exact section (no retrieval)1.6% (5/307)30795% CI 0.7% to 3.8%2026-08-24
Statute pinpoint, exact section (with DocketBuster retrieval)75.6% (232/307)30795% CI 70.5% to 80.0%2026-08-24
Say-nothing rate (declines to bluff when the answer is not in the record)100.0%60095% CI 99.4% to 100.0%2026-08-25
Abstained on statute pinpoint17.9% (55/307)307count, no interval reported2026-08-24
Coaching quality (GW-14x, 0 to 8)3.56 / 8-rubric mean, no interval reported2026-08-24

Source files: hard-llm-level-nemotron3-nano-30b-statute.json, hard-llm-level-nemotron3-nano-30b-statute_rag.json, gw14x-ortier-nemotron3nano.json, hard-llm-level-nemotron3-nano-30b-saynothing.json. Raw model name in source: nvidia/nemotron-3-nano-30b-a3b.