Verify the benchmark
Scores are only worth something if we cannot quietly change the questions after seeing how a model did. Every question in the bank and every graded answer is hashed into a Merkle tree. We publish the roots. Recompute them yourself from the published data, or hold us to a root we published before a run.
Runs
| Run | Grader | Cases | Mean | Results root |
|---|---|---|---|---|
| deepseek/deepseek-v4-flash Raw (model alone) | hll-grader-0.5 | 2001 | 39% | 87a6e5ff1164882c38b0c2d0… |
| deepseek/deepseek-v4-flash With DocketRouter | hll-grader-0.5 | 1175 | 51% | 9b57440016a098b12c341218… |
Grader source
The scoring code is mechanical and hashed with every attestation, so a score can always be tied to the exact code that produced it. One digest per scoring module:
How to check it yourself
Recompute: canonical-JSON each item (keys sorted at every depth), leaf = sha256(0x00||json), sort leaves by item id, then fold pairs with sha256(0x01||left||right). The result must equal bank.merkle_root. For a run: leaf = sha256(0x00||canon({item, answer_sha256, score})), sort leaves, fold the same way.
The public split is downloadable at /api/v1/hll/public. A private question can be proven to have existed unchanged by revealing that one question plus its Merkle path, without publishing the rest of the private split.