Solidity review
-
GPT-6 Luna
openai/gpt-6-luna · raw arm
- 909 of 10
- $0.011
-
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash with the Solidity review profile
- 909 of 10
- $0.05
Paydirt · release 2026-10
Every model runs through one agent harness at its highest reasoning effort on 18 held pairs. A pair counts only when the model is right on both twins.
Ranked by score and grouped by tier, in the order of the results file.
| # | Model | Effort sent | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Tier 1intervals overlap tier leader | |||||||||
| 1 | GPT-6 Astraopenai/gpt-6-astraopenai | Score94.417 of 18 pairs95% interval 83.3 to 100 | Recall100% | Fool’s gold8% | Challenge100% | Failures0% | Cost / run$0.52 | Time2m 21s | Effortmax |
| 1 | GPT-6.1 Solopenai/gpt-6.1-solopenai | Score94.417 of 18 pairs95% interval 83.3 to 100 | Recall100% | Fool’s gold8% | Challenge100% | Failures0% | Cost / run$0.108 | Time2m 34s | Effortmax |
| 3 | Kimi K3moonshotai/kimi-k3moonshotaiOpen weight | Score88.916 of 18 pairs95% interval 72.2 to 100 | Recall83% | Fool’s gold0% | Challenge100% | Failures0% | Cost / run$0.22 | Time4m 02s | Effortmax |
| 3 | GPT-5.6 Terraopenai/gpt-5.6-terraopenai | Score88.916 of 18 pairs95% interval 72.2 to 100 | Recall92% | Fool’s gold8% | Challenge100% | Failures0% | Cost / run$0.433 | Time5m 09s | Effortmax |
| 5 | Claude Opus 5.5anthropic/claude-opus-5.5anthropic | Score83.315 of 18 pairs95% interval 66.7 to 100 | Recall83% | Fool’s gold8% | Challenge100% | Failures11% | Cost / run$0.882 | Time5m 58s | Effortmax |
| 5 | GPT-6 Lunaopenai/gpt-6-lunaopenai | Score83.315 of 18 pairs95% interval 66.7 to 100 | Recall92% | Fool’s gold8% | Challenge92% | Failures0% | Cost / run$0.012 | Time2m 40s | Effortmax |
| 7 | Claude Fable 5.1anthropic/claude-fable-5.1anthropic | Score77.814 of 18 pairs95% interval 55.6 to 94.4 | Recall83% | Fool’s gold8% | Challenge92% | Failures8% | Cost / run$1.38 | Time4m 17s | Effortmax |
| 7 | DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813deepseekOpen weight | Score77.814 of 18 pairs95% interval 55.6 to 94.4 | Recall75% | Fool’s gold8% | Challenge100% | Failures0% | Cost / run$0.088 | Time2m 00s | Effortmax |
| 7 | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight | Score77.814 of 18 pairs95% interval 55.6 to 94.4 | Recall92% | Fool’s gold8% | Challenge100% | Failures6% | Cost / run$0.063 | Time5m 13s | Effortmax |
| 7 | Hy4 previewtencent/hy4-previewtencentOpen weight | Score77.814 of 18 pairs95% interval 55.6 to 94.4 | Recall75% | Fool’s gold8% | Challenge100% | Failures0% | Cost / run$0.087 | Time4m 38s | Efforthighasked max |
| 7 | GLM 5.3z-ai/glm-5.3z-aiOpen weight | Score77.814 of 18 pairs95% interval 55.6 to 94.4 | Recall83% | Fool’s gold17% | Challenge100% | Failures0% | Cost / run$0.105 | Time5m 18s | Effortmax |
| 12 | Qwen3.8 27Bqwen/qwen3.8-27bqwenOpen weight | Score72.213 of 18 pairs95% interval 50 to 94.4 | Recall58% | Fool’s gold0% | Challenge100% | Failures0% | Cost / run$0.085 | Time2m 58s | Efforthighasked maxmixed: high 28, xhigh 8 |
| 12 | GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight | Score72.213 of 18 pairs95% interval 50 to 88.9 | Recall75% | Fool’s gold0% | Challenge83% | Failures6% | Cost / run$0.016 | Time6m 09s | Effortmax |
| 14 | Qwen3.8 Max (0902)qwen/qwen3.8-max-0902qwen | Score66.712 of 18 pairs95% interval 44.4 to 88.9 | Recall75% | Fool’s gold25% | Challenge100% | Failures0% | Cost / run$0.169 | Time5m 05s | Effortxhighasked max |
| 15 | MiniMax M3minimax/minimax-m3minimaxOpen weight | Score61.111 of 18 pairs95% interval 38.9 to 83.3 | Recall58% | Fool’s gold17% | Challenge100% | Failures8% | Cost / run$0.039 | Time2m 51s | Efforthighasked max |
| Tier 2intervals overlap tier leader | |||||||||
| 16 | Gemini 3.8 Flashgoogle/gemini-3.8-flashgoogle | Score55.610 of 18 pairs95% interval 33.3 to 77.8 | Recall58% | Fool’s gold17% | Challenge100% | Failures22% | Cost / run$0.204 | Time2m 55s | Efforthighasked max |
| 16 | Grok 4.7x-ai/grok-4.7x-ai | Score55.610 of 18 pairs95% interval 33.3 to 77.8 | Recall67% | Fool’s gold17% | Challenge75% | Failures31% | Cost / run$0.373 | Time12m 10s | Effortxhighasked max |
| 18 | Claude Haiku 4.5anthropic/claude-haiku-4.5anthropic | Score44.48 of 18 pairs95% interval 22.2 to 66.7 | Recall33% | Fool’s gold0% | Challenge83% | Failures0% | Cost / run$0.165 | Time47s | Efforthighasked max |
| 18 | Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55bnvidiaOpen weight | Score44.48 of 18 pairs95% interval 22.2 to 66.7 | Recall33% | Fool’s gold17% | Challenge92% | Failures3% | Cost / run$0.045 | Time5m 08s | Efforthighasked max |
| 20 | Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-previewgoogle | Score38.97 of 18 pairs95% interval 16.7 to 61.1 | Recall83% | Fool’s gold0% | Challenge100% | Failures36% | Cost / run$0.147 | Time8m 09s | Efforthighasked max |
| 21 | Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic | Score33.36 of 18 pairs95% interval 11.1 to 55.6 | Recall83% | Fool’s gold0% | Challenge83% | Failures36% | Cost / run$1.09 | Time12m 01s | Effortmax |
| 21 | Mistral Medium 3.5mistralai/mistral-medium-3-5mistralai | Score33.36 of 18 pairs95% interval 11.1 to 55.6 | Recall33% | Fool’s gold42% | Challenge83% | Failures3% | Cost / run$0.174 | Time1m 45s | Efforthighasked max |
A tier starts at its highest-scoring model and takes in every model whose 95% interval overlaps that model’s. These are descriptive groups; overlapping intervals do not establish equal performance.
The highest score on each kind of task, the best of the models under $0.10 a run, and the best open-weight model. Equal scores go to the cheaper run.
GPT-6 Luna
openai/gpt-6-luna · raw arm
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash with the Solidity review profile
GLM 5.3 Flash
z-ai/glm-5.3-flash · raw arm
GLM 5.3 Flash
z-ai/glm-5.3-flash with the Challenge a draft report profile
Scores count only the pairs of that task. Costs are medians across every input in the selected arm, which can include other tasks. The workbench applies the matching core profile; raw scores do not measure that profile. Pick the model there.
The same model on the same pairs, once raw and once with the matching Bounty Operator core profile. A profile arm differs from the raw arm in one thing: the profile’s system message is appended to the system prompt. The task and the answer sheet are the same. Profile arms ran for four models, chosen before any scored run.
| Model | Profile | Raw | With profile | Lift | 95% interval | Significant |
|---|---|---|---|---|---|---|
| GPT-6 Lunaopenai/gpt-6-lunaopenai | ProfileChallenge a draft report6 draft-report pairs | Raw83.3 | With profile50 | Lift−33.3 | 95% interval−83.3 to +33.3 | SignificantNo |
| GPT-6 Lunaopenai/gpt-6-lunaopenai | ProfileSolidity review10 Solidity pairs | Raw90 | With profile80 | Lift−10 | 95% interval−30 to 0 | SignificantNo |
| GPT-6 Lunaopenai/gpt-6-lunaopenai | ProfileCode security review2 TypeScript API pairs | Raw50 | With profile50 | Lift0 | 95% interval0 to 0 | SignificantNo |
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight | ProfileChallenge a draft report6 draft-report pairs | Raw100 | With profile100 | Lift0 | 95% interval0 to 0 | SignificantNo |
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight | ProfileSolidity review10 Solidity pairs | Raw70 | With profile90 | Lift+20 | 95% interval0 to +50 | SignificantNo |
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight | ProfileCode security review2 TypeScript API pairs | Raw50 | With profile100 | Lift+50 | 95% interval0 to +100 | SignificantNo |
| GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight | ProfileChallenge a draft report6 draft-report pairs | Raw66.7 | With profile100 | Lift+33.3 | 95% interval0 to +66.7 | SignificantNo |
| GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight | ProfileSolidity review10 Solidity pairs | Raw70 | With profile60 | Lift−10 | 95% interval−40 to +20 | SignificantNo |
| GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight | ProfileCode security review2 TypeScript API pairs | Raw100 | With profile100 | Lift0 | 95% interval0 to 0 | SignificantNo |
| Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic | ProfileChallenge a draft report6 draft-report pairs | Raw66.7 | With profile0 | Lift−66.7 | 95% interval−100 to −33.3 | SignificantYes |
| Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic | ProfileSolidity review10 Solidity pairs | Raw10 | With profile0 | Lift−10 | 95% interval−30 to 0 | SignificantNo |
| Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic | ProfileCode security review2 TypeScript API pairs | Raw50 | With profile0 | Lift−50 | 95% interval−100 to 0 | SignificantNo |
Lift is the profile score minus the raw score on the same pairs, in points, with a paired bootstrap interval. It is called significant only when the interval excludes zero. Negative lifts are published as they are.
Every ranked model on each of the 18 held pairs, on the raw arm. Right means right on both inputs of the pair. Failed means at least one input ended with no readable answer.
| Solidity | TypeScript API | Draft report | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | hard-01 | hard-02 | hard-03 | hard-04 | sol-02 | sol-03 | sol-05 | sol-09 | sol-10 | sol-11 | ts-02 | ts-03 | ch-02 | ch-05 | ch-06 | ch-07 | hard-07 | hard-08 | Right |
| GPT-6 Astra | hard-01: right | hard-02: wrong | hard-03: right | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 17 |
| GPT-6.1 Sol | hard-01: right | hard-02: wrong | hard-03: right | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 17 |
| Kimi K3 | hard-01: right | hard-02: right | hard-03: wrong | hard-04: wrong | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 16 |
| GPT-5.6 Terra | hard-01: right | hard-02: right | hard-03: wrong | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 16 |
| Claude Opus 5.5 | hard-01: right | hard-02: failed | hard-03: right | hard-04: failed | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: failed | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 15 |
| GPT-6 Luna | hard-01: right | hard-02: right | hard-03: wrong | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: wrong | hard-07: right | hard-08: right | 15 |
| Claude Fable 5.1 | hard-01: right | hard-02: right | hard-03: failed | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: failed | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: failed | ch-07: right | hard-07: right | hard-08: right | 14 |
| DeepSeek V4 Pro 0813 | hard-01: wrong | hard-02: right | hard-03: wrong | hard-04: wrong | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 14 |
| DeepSeek V4.1 Flash | hard-01: right | hard-02: failed | hard-03: wrong | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: failed | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 14 |
| Hy4 preview | hard-01: wrong | hard-02: wrong | hard-03: wrong | hard-04: wrong | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 14 |
| GLM 5.3 | hard-01: right | hard-02: right | hard-03: wrong | hard-04: wrong | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: wrong | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 14 |
| Qwen3.8 27B | hard-01: right | hard-02: right | hard-03: wrong | hard-04: wrong | sol-02: right | sol-03: right | sol-05: right | sol-09: wrong | sol-10: right | sol-11: wrong | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 13 |
| GLM 5.3 Flash | hard-01: right | hard-02: wrong | hard-03: wrong | hard-04: right | sol-02: right | sol-03: right | sol-05: right | sol-09: right | sol-10: right | sol-11: wrong | ts-02: right | ts-03: right | ch-02: right | ch-05: failed | ch-06: right | ch-07: right | hard-07: right | hard-08: failed | 13 |
| Qwen3.8 Max (0902) | hard-01: wrong | hard-02: right | hard-03: wrong | hard-04: wrong | sol-02: right | sol-03: wrong | sol-05: right | sol-09: wrong | sol-10: right | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 12 |
| MiniMax M3 | hard-01: wrong | hard-02: wrong | hard-03: wrong | hard-04: failed | sol-02: wrong | sol-03: failed | sol-05: right | sol-09: wrong | sol-10: right | sol-11: right | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 11 |
| Gemini 3.8 Flash | hard-01: failed | hard-02: failed | hard-03: wrong | hard-04: failed | sol-02: failed | sol-03: failed | sol-05: failed | sol-09: failed | sol-10: right | sol-11: right | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 10 |
| Grok 4.7 | hard-01: failed | hard-02: failed | hard-03: failed | hard-04: failed | sol-02: right | sol-03: wrong | sol-05: right | sol-09: right | sol-10: right | sol-11: right | ts-02: right | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: failed | hard-07: failed | hard-08: right | 10 |
| Claude Haiku 4.5 | hard-01: wrong | hard-02: wrong | hard-03: wrong | hard-04: wrong | sol-02: wrong | sol-03: right | sol-05: right | sol-09: wrong | sol-10: wrong | sol-11: wrong | ts-02: right | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: wrong | hard-08: wrong | 8 |
| Nemotron 3 Ultra | hard-01: right | hard-02: wrong | hard-03: wrong | hard-04: wrong | sol-02: wrong | sol-03: wrong | sol-05: right | sol-09: failed | sol-10: wrong | sol-11: wrong | ts-02: wrong | ts-03: right | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: wrong | hard-08: right | 8 |
| Gemini 3.1 Pro Preview | hard-01: failed | hard-02: failed | hard-03: failed | hard-04: failed | sol-02: failed | sol-03: failed | sol-05: failed | sol-09: failed | sol-10: right | sol-11: failed | ts-02: failed | ts-03: failed | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: right | hard-08: right | 7 |
| Claude Sonnet 5.5 | hard-01: failed | hard-02: failed | hard-03: failed | hard-04: failed | sol-02: failed | sol-03: failed | sol-05: right | sol-09: failed | sol-10: failed | sol-11: failed | ts-02: right | ts-03: failed | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: failed | hard-08: failed | 6 |
| Mistral Medium 3.5 | hard-01: wrong | hard-02: wrong | hard-03: wrong | hard-04: wrong | sol-02: wrong | sol-03: wrong | sol-05: right | sol-09: wrong | sol-10: right | sol-11: wrong | ts-02: wrong | ts-03: wrong | ch-02: right | ch-05: right | ch-06: right | ch-07: right | hard-07: failed | hard-08: wrong | 6 |
Pair names only: the held pairs are not published. Bug classes of the scored find pairs (the label of a held pair is not published, the distribution is): language semantics 2, oracle staleness and decimals 2, signature replay and domain separation 2, access control on sibling functions 1, cross-chain message replay 1, initialisation and upgrade gaps 1, queue and epoch boundary errors 1, reward checkpoint order 1, share inflation and rounding direction 1. No class appears more than 2 times.
Each claim, and the mechanism behind it.
Every case was written for Paydirt. 36 held cases carry a salted SHA-256 commitment written before the first scored run; a case’s salt is published when it moves to the public set.
Each pair is a vulnerable twin and its fix, or an overclaimed draft and an accurate one. Flagging everything, or nothing, scores zero for the pair.
A finding hits when it names the right file and either the right function or overlapping lines. A challenge counts by verdict and a quote found at a false statement of the draft. No model judges a model.
Every run is a fresh omp 18.4.4 process with a throwaway home, a fresh copy of the workspace and three read-only tools: read, grep and glob.
The harness asks every model for “max”, its highest effort, and records the effort each request carried. The table shows it per model.
Every request carries the same provider-routing block: no endpoint that declares less than 8-bit precision, slow endpoints last. A rate limit, a server error or a dropped connection is retried, never scored. The latest stored invocations record 29 automatic retries; earlier invocations and interrupted records are excluded from this count.
| What | SHA-256 |
|---|---|
| Protocol | 74aad5311600bf53862cb0c5e09dec6e6af0f63feaff7d6805d021905e0cfd35 |
Arm raw (the core) | 5fea9eb00e3f2bb24d3965bf6854d39742d3324b1b826492eaa1ef51dfe09768 |
Arm solidity | ad6a658a538ec5ef84c3025febf9dea8a5e076127e2b11e9dcfe346cb2d38a15 |
Arm general | d7f99fc3799b5bf3914fd3aba5668fdf9467a391a1607f896c0e9d26083a01fc |
Arm report | 4550f8a9890f6efd39f757fcea9c5b7ea870a7beccf9fc84c715d66c7a7f6448 |
Download latest.json, the per-model files and the archive, unpack the archive and run these from its folder.
Needs Node 22 or later and nothing else. It recomputes every score, interval, tier and lift from the per-run outcomes, recomputes the hashes from the files in the archive, and re-scores every run on a public case from its stored output.
node bench/bench.mjs verify --results latest.json --archive paydirt-2026-10-public.tar.gz
node bench/bench.mjs verify --results practice/latest.json --archive paydirt-2026-10-public.tar.gz
Needs Foundry for the Solidity proofs. Every planted bug must pass its proof on the vulnerable twin and fail it on the fixed one.
node bench/bench.mjs lint --proofs
Needs omp 18.4.4 and an OpenRouter key. Models are not deterministic: a rerun lands near the published numbers, not on them.
node bench/bench.mjs run --set public --models <slug> --run-id mine
node bench/bench.mjs score --set public --run-id mine
6 public pairs, published with their answer keys, proofs and the raw output of every run on them, so every outcome here can be re-scored end to end. This is not the Paydirt Score.
| Model | Pairs right | Recall | Fool’s gold | Cost per run |
|---|---|---|---|---|
| GPT-6 Luna | Pairs right6 of 6 | Recall100% | Fool’s gold0% | Cost per run$0.0086 |
| DeepSeek V4.1 Flash | Pairs right5 of 6 | Recall100% | Fool’s gold0% | Cost per run$0.02 |
| GLM 5.3 Flash | Pairs right4 of 6 | Recall100% | Fool’s gold0% | Cost per run$0.011 |
3 models of the release’s list have no score. A model with no answer is listed here, not given a low score.
| Why | Models |
|---|---|
| Its hosts returned tool calls as plain text on 35 of 36 initial inputs. The batch could not be completed and is excluded from the ranking. | Model
|
| OpenRouter serves it only to accounts that have made an age confirmation, which the benchmark account has not made. | Model
|
| Its runs were not finished when this release was published: 35 of 36 inputs completed. | Model
|
Google Gemini 3.5 Flash Lite has 35 finalized raw inputs and one unresolved input after repeated provider failures. Its results are shown for inspection, but it is not ranked or recommended.
Each pair is an original case with an executable proof for every planted bug and every decoy, and the scored pairs are held, so the set grows slowly. With 18 pairs the 95% interval is wide; it is published so that a two-pair difference is not read as a ranking. Every model answers every input once. The release’s credit covered one repeat for many models rather than three for a handful, so the score carries the noise of one run per input: the interval covers which pairs were drawn, not how a model varies between runs.
Models run in tiers, the cheapest first inside a tier, until the release’s credit is used. A model the credit did not reach, or whose host returned no answer, is listed under Not run with the reason. A model whose runs had not finished at publication is listed there too, with how many of its inputs were completed. None of them is given a low score.
Bounty Operator is not a row. The ranked score comes from the raw arm, which sends no product text: the same system prompt, task and answer sheet to every model. The profile arms, which add a Bounty Operator core profile, are published separately, negative lifts included. Each pair started as a draft by a model of one of nine vendors. Sessions of Claude Opus 5.5, a ranked model, directed by the benchmark’s owner, then checked, repaired and proved every pair and wrote two of the hard pairs. Those two count as Anthropic-drafted, and the results file gives every model’s score without the pairs its own vendor drafted. Two ranked models, DeepSeek V4.1 Flash and GLM 5.3 Flash, ran the pilot on the find pairs of that time, and while nine pairs were written, drafts were sent to a few ranked models to see whether they could be decided both ways; the method names the models and the pairs whose design changed after such a check.
A held case’s files are not published. Before the first scored run each held case got a salted SHA-256 commitment (36 in this release); when a held case moves to the public set, its salt is published with it so the commitment can be checked. Held does not mean unseen by the model providers: a held case is sent to the providers that serve each model, as every prompt is.
With each release. A release spends a fixed amount of credit in run order: tier 1 first, the cheapest model first inside a tier. Held pairs move to the public set over time and new held pairs replace them. A change to the harness, the routing policy, the task text or a set changes every hash, and no earlier run counts after it.