Paydirt · release 2026-10

Which model finds the real bug, leaves the fixed code alone and catches an overclaimed report

Every model runs through one agent harness at its highest reasoning effort on 18 held pairs. A pair counts only when the model is right on both twins.

Models ranked
22
Held pairs
18
Inputs per model
36
Recorded inference cost
$273.01
Published
3 October 2026
Harness
omp 18.4.4

The leaderboard

Ranked by score and grouped by tier, in the order of the results file.

Paydirt release 2026-10: 22 ranked models on 18 held pairs, in the order of the results file, grouped by tier.
#ModelEffort sent
Tier 1intervals overlap tier leader
1GPT-6 Astraopenai/gpt-6-astraopenaiScore94.417 of 18 pairs95% interval 83.3 to 100Recall100%Fool’s gold8%Challenge100%Failures0%Cost / run$0.52Time2m 21sEffortmax
1GPT-6.1 Solopenai/gpt-6.1-solopenaiScore94.417 of 18 pairs95% interval 83.3 to 100Recall100%Fool’s gold8%Challenge100%Failures0%Cost / run$0.108Time2m 34sEffortmax
3Kimi K3moonshotai/kimi-k3moonshotaiOpen weightScore88.916 of 18 pairs95% interval 72.2 to 100Recall83%Fool’s gold0%Challenge100%Failures0%Cost / run$0.22Time4m 02sEffortmax
3GPT-5.6 Terraopenai/gpt-5.6-terraopenaiScore88.916 of 18 pairs95% interval 72.2 to 100Recall92%Fool’s gold8%Challenge100%Failures0%Cost / run$0.433Time5m 09sEffortmax
5Claude Opus 5.5anthropic/claude-opus-5.5anthropicScore83.315 of 18 pairs95% interval 66.7 to 100Recall83%Fool’s gold8%Challenge100%Failures11%Cost / run$0.882Time5m 58sEffortmax
5GPT-6 Lunaopenai/gpt-6-lunaopenaiScore83.315 of 18 pairs95% interval 66.7 to 100Recall92%Fool’s gold8%Challenge92%Failures0%Cost / run$0.012Time2m 40sEffortmax
7Claude Fable 5.1anthropic/claude-fable-5.1anthropicScore77.814 of 18 pairs95% interval 55.6 to 94.4Recall83%Fool’s gold8%Challenge92%Failures8%Cost / run$1.38Time4m 17sEffortmax
7DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813deepseekOpen weightScore77.814 of 18 pairs95% interval 55.6 to 94.4Recall75%Fool’s gold8%Challenge100%Failures0%Cost / run$0.088Time2m 00sEffortmax
7DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weightScore77.814 of 18 pairs95% interval 55.6 to 94.4Recall92%Fool’s gold8%Challenge100%Failures6%Cost / run$0.063Time5m 13sEffortmax
7Hy4 previewtencent/hy4-previewtencentOpen weightScore77.814 of 18 pairs95% interval 55.6 to 94.4Recall75%Fool’s gold8%Challenge100%Failures0%Cost / run$0.087Time4m 38sEfforthighasked max
7GLM 5.3z-ai/glm-5.3z-aiOpen weightScore77.814 of 18 pairs95% interval 55.6 to 94.4Recall83%Fool’s gold17%Challenge100%Failures0%Cost / run$0.105Time5m 18sEffortmax
12Qwen3.8 27Bqwen/qwen3.8-27bqwenOpen weightScore72.213 of 18 pairs95% interval 50 to 94.4Recall58%Fool’s gold0%Challenge100%Failures0%Cost / run$0.085Time2m 58sEfforthighasked maxmixed: high 28, xhigh 8
12GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weightScore72.213 of 18 pairs95% interval 50 to 88.9Recall75%Fool’s gold0%Challenge83%Failures6%Cost / run$0.016Time6m 09sEffortmax
14Qwen3.8 Max (0902)qwen/qwen3.8-max-0902qwenScore66.712 of 18 pairs95% interval 44.4 to 88.9Recall75%Fool’s gold25%Challenge100%Failures0%Cost / run$0.169Time5m 05sEffortxhighasked max
15MiniMax M3minimax/minimax-m3minimaxOpen weightScore61.111 of 18 pairs95% interval 38.9 to 83.3Recall58%Fool’s gold17%Challenge100%Failures8%Cost / run$0.039Time2m 51sEfforthighasked max
Tier 2intervals overlap tier leader
16Gemini 3.8 Flashgoogle/gemini-3.8-flashgoogleScore55.610 of 18 pairs95% interval 33.3 to 77.8Recall58%Fool’s gold17%Challenge100%Failures22%Cost / run$0.204Time2m 55sEfforthighasked max
16Grok 4.7x-ai/grok-4.7x-aiScore55.610 of 18 pairs95% interval 33.3 to 77.8Recall67%Fool’s gold17%Challenge75%Failures31%Cost / run$0.373Time12m 10sEffortxhighasked max
18Claude Haiku 4.5anthropic/claude-haiku-4.5anthropicScore44.48 of 18 pairs95% interval 22.2 to 66.7Recall33%Fool’s gold0%Challenge83%Failures0%Cost / run$0.165Time47sEfforthighasked max
18Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55bnvidiaOpen weightScore44.48 of 18 pairs95% interval 22.2 to 66.7Recall33%Fool’s gold17%Challenge92%Failures3%Cost / run$0.045Time5m 08sEfforthighasked max
20Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-previewgoogleScore38.97 of 18 pairs95% interval 16.7 to 61.1Recall83%Fool’s gold0%Challenge100%Failures36%Cost / run$0.147Time8m 09sEfforthighasked max
21Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropicScore33.36 of 18 pairs95% interval 11.1 to 55.6Recall83%Fool’s gold0%Challenge83%Failures36%Cost / run$1.09Time12m 01sEffortmax
21Mistral Medium 3.5mistralai/mistral-medium-3-5mistralaiScore33.36 of 18 pairs95% interval 11.1 to 55.6Recall33%Fool’s gold42%Challenge83%Failures3%Cost / run$0.174Time1m 45sEfforthighasked max

A tier starts at its highest-scoring model and takes in every model whose 95% interval overlaps that model’s. These are descriptive groups; overlapping intervals do not establish equal performance.

What each column means
Paydirt Score
100 × pairs right ÷ 18, on the raw arm. A pair is right only when both of its inputs are right. The bar is the 95% interval.
Recall
Planted bugs hit ÷ planted bugs, on the vulnerable twins.
Fool’s gold
Fixed twins whose patched code was reported as a bug rated medium, high, critical or unrated ÷ fixed twins.
Challenge accuracy
The mean of two shares: overclaimed drafts caught with a quote of a false statement, and accurate drafts accepted.
Failure rate
Inputs that ended with no readable answer (refusal, timeout, truncation, a provider error of the model’s own, an answer sheet that does not parse) ÷ inputs.
Cost per run
The mean recorded cost per input in US dollars, including preserved retries. Unpreserved attempts are excluded, so the recorded cost can be lower than total account spend.
Median time
Wall time of the median run.
Effort sent
The reasoning effort the recorded requests carried. The harness asks every model for its highest.

What the numbers say

  1. 15 models share the top tier, from 61.1 to 94.4: each one's 95% interval overlaps the leader's.
  2. In that tier, GPT-6 Luna costs $0.012 a run and Claude Fable 5.1 $1.38, 118.5 times as much.
  3. Across the 22 ranked models, a run costs from $0.012 (GPT-6 Luna) to $1.38 (Claude Fable 5.1).
  4. Eleven of the 22 ranked models returned a readable answer on all 36 inputs.

Which model to use for what

The highest score on each kind of task, the best of the models under $0.10 a run, and the best open-weight model. Equal scores go to the cheaper run.

Solidity review

10 pairs
  • Best scoreUnder $0.10 a run

    GPT-6 Luna

    openai/gpt-6-luna · raw arm

    Score
    909 of 10
    Median cost across arm
    $0.011
    Open the workbench
  • Open weight

    DeepSeek V4.1 Flash

    deepseek/deepseek-v4.1-flash with the Solidity review profile

    Score
    909 of 10
    Median cost across arm
    $0.05
    Open the workbench

TypeScript API review

2 pairs
  • Best scoreUnder $0.10 a runOpen weight

    GLM 5.3 Flash

    z-ai/glm-5.3-flash · raw arm

    Score
    1002 of 2
    Median cost across arm
    $0.014
    Open the workbench

Challenging a draft report

6 pairs
  • Best scoreUnder $0.10 a runOpen weight

    GLM 5.3 Flash

    z-ai/glm-5.3-flash with the Challenge a draft report profile

    Score
    1006 of 6
    Median cost across arm
    $0.018
    Open the workbench

Scores count only the pairs of that task. Costs are medians across every input in the selected arm, which can include other tasks. The workbench applies the matching core profile; raw scores do not measure that profile. Pick the model there.

Profile lift

The same model on the same pairs, once raw and once with the matching Bounty Operator core profile. A profile arm differs from the raw arm in one thing: the profile’s system message is appended to the system prompt. The task and the answer sheet are the same. Profile arms ran for four models, chosen before any scored run.

ModelProfileRawWith profileLift95% intervalSignificant
GPT-6 Lunaopenai/gpt-6-lunaopenai ProfileChallenge a draft report6 draft-report pairs Raw83.3 With profile50 Lift−33.3 95% interval−83.3 to +33.3 SignificantNo
GPT-6 Lunaopenai/gpt-6-lunaopenai ProfileSolidity review10 Solidity pairs Raw90 With profile80 Lift−10 95% interval−30 to 0 SignificantNo
GPT-6 Lunaopenai/gpt-6-lunaopenai ProfileCode security review2 TypeScript API pairs Raw50 With profile50 Lift0 95% interval0 to 0 SignificantNo
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight ProfileChallenge a draft report6 draft-report pairs Raw100 With profile100 Lift0 95% interval0 to 0 SignificantNo
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight ProfileSolidity review10 Solidity pairs Raw70 With profile90 Lift+20 95% interval0 to +50 SignificantNo
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashdeepseekOpen weight ProfileCode security review2 TypeScript API pairs Raw50 With profile100 Lift+50 95% interval0 to +100 SignificantNo
GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight ProfileChallenge a draft report6 draft-report pairs Raw66.7 With profile100 Lift+33.3 95% interval0 to +66.7 SignificantNo
GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight ProfileSolidity review10 Solidity pairs Raw70 With profile60 Lift−10 95% interval−40 to +20 SignificantNo
GLM 5.3 Flashz-ai/glm-5.3-flashz-aiOpen weight ProfileCode security review2 TypeScript API pairs Raw100 With profile100 Lift0 95% interval0 to 0 SignificantNo
Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic ProfileChallenge a draft report6 draft-report pairs Raw66.7 With profile0 Lift−66.7 95% interval−100 to −33.3 SignificantYes
Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic ProfileSolidity review10 Solidity pairs Raw10 With profile0 Lift−10 95% interval−30 to 0 SignificantNo
Claude Sonnet 5.5anthropic/claude-sonnet-5.5anthropic ProfileCode security review2 TypeScript API pairs Raw50 With profile0 Lift−50 95% interval−100 to 0 SignificantNo

Lift is the profile score minus the raw score on the same pairs, in points, with a paired bootstrap interval. It is called significant only when the interval excludes zero. Negative lifts are published as they are.

Pair by pair

Every ranked model on each of the 18 held pairs, on the raw arm. Right means right on both inputs of the pair. Failed means at least one input ended with no readable answer.

SolidityTypeScript APIDraft report
Modelhard-01hard-02hard-03hard-04sol-02sol-03sol-05sol-09sol-10sol-11ts-02ts-03ch-02ch-05ch-06ch-07hard-07hard-08Right
GPT-6 Astrahard-01: righthard-02: wronghard-03: righthard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right17
GPT-6.1 Solhard-01: righthard-02: wronghard-03: righthard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right17
Kimi K3hard-01: righthard-02: righthard-03: wronghard-04: wrongsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right16
GPT-5.6 Terrahard-01: righthard-02: righthard-03: wronghard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right16
Claude Opus 5.5hard-01: righthard-02: failedhard-03: righthard-04: failedsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: failedch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right15
GPT-6 Lunahard-01: righthard-02: righthard-03: wronghard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: wronghard-07: righthard-08: right15
Claude Fable 5.1hard-01: righthard-02: righthard-03: failedhard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: failedsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: failedch-07: righthard-07: righthard-08: right14
DeepSeek V4 Pro 0813hard-01: wronghard-02: righthard-03: wronghard-04: wrongsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right14
DeepSeek V4.1 Flashhard-01: righthard-02: failedhard-03: wronghard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: failedts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right14
Hy4 previewhard-01: wronghard-02: wronghard-03: wronghard-04: wrongsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right14
GLM 5.3hard-01: righthard-02: righthard-03: wronghard-04: wrongsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: wrongsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right14
Qwen3.8 27Bhard-01: righthard-02: righthard-03: wronghard-04: wrongsol-02: rightsol-03: rightsol-05: rightsol-09: wrongsol-10: rightsol-11: wrongts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right13
GLM 5.3 Flashhard-01: righthard-02: wronghard-03: wronghard-04: rightsol-02: rightsol-03: rightsol-05: rightsol-09: rightsol-10: rightsol-11: wrongts-02: rightts-03: rightch-02: rightch-05: failedch-06: rightch-07: righthard-07: righthard-08: failed13
Qwen3.8 Max (0902)hard-01: wronghard-02: righthard-03: wronghard-04: wrongsol-02: rightsol-03: wrongsol-05: rightsol-09: wrongsol-10: rightsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right12
MiniMax M3hard-01: wronghard-02: wronghard-03: wronghard-04: failedsol-02: wrongsol-03: failedsol-05: rightsol-09: wrongsol-10: rightsol-11: rightts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right11
Gemini 3.8 Flashhard-01: failedhard-02: failedhard-03: wronghard-04: failedsol-02: failedsol-03: failedsol-05: failedsol-09: failedsol-10: rightsol-11: rightts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right10
Grok 4.7hard-01: failedhard-02: failedhard-03: failedhard-04: failedsol-02: rightsol-03: wrongsol-05: rightsol-09: rightsol-10: rightsol-11: rightts-02: rightts-03: wrongch-02: rightch-05: rightch-06: rightch-07: failedhard-07: failedhard-08: right10
Claude Haiku 4.5hard-01: wronghard-02: wronghard-03: wronghard-04: wrongsol-02: wrongsol-03: rightsol-05: rightsol-09: wrongsol-10: wrongsol-11: wrongts-02: rightts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: wronghard-08: wrong8
Nemotron 3 Ultrahard-01: righthard-02: wronghard-03: wronghard-04: wrongsol-02: wrongsol-03: wrongsol-05: rightsol-09: failedsol-10: wrongsol-11: wrongts-02: wrongts-03: rightch-02: rightch-05: rightch-06: rightch-07: righthard-07: wronghard-08: right8
Gemini 3.1 Pro Previewhard-01: failedhard-02: failedhard-03: failedhard-04: failedsol-02: failedsol-03: failedsol-05: failedsol-09: failedsol-10: rightsol-11: failedts-02: failedts-03: failedch-02: rightch-05: rightch-06: rightch-07: righthard-07: righthard-08: right7
Claude Sonnet 5.5hard-01: failedhard-02: failedhard-03: failedhard-04: failedsol-02: failedsol-03: failedsol-05: rightsol-09: failedsol-10: failedsol-11: failedts-02: rightts-03: failedch-02: rightch-05: rightch-06: rightch-07: righthard-07: failedhard-08: failed6
Mistral Medium 3.5hard-01: wronghard-02: wronghard-03: wronghard-04: wrongsol-02: wrongsol-03: wrongsol-05: rightsol-09: wrongsol-10: rightsol-11: wrongts-02: wrongts-03: wrongch-02: rightch-05: rightch-06: rightch-07: righthard-07: failedhard-08: wrong6

Pair names only: the held pairs are not published. Bug classes of the scored find pairs (the label of a held pair is not published, the distribution is): language semantics 2, oracle staleness and decimals 2, signature replay and domain separation 2, access control on sibling functions 1, cross-chain message replay 1, initialisation and upgrade gaps 1, queue and epoch boundary errors 1, reward checkpoint order 1, share inflation and rounding direction 1. No class appears more than 2 times.

How it is kept honest

Each claim, and the mechanism behind it.

Hashes this release ran under

WhatSHA-256
Protocol74aad5311600bf53862cb0c5e09dec6e6af0f63feaff7d6805d021905e0cfd35
Arm raw (the core)5fea9eb00e3f2bb24d3965bf6854d39742d3324b1b826492eaa1ef51dfe09768
Arm solidityad6a658a538ec5ef84c3025febf9dea8a5e076127e2b11e9dcfe346cb2d38a15
Arm generald7f99fc3799b5bf3914fd3aba5668fdf9467a391a1607f896c0e9d26083a01fc
Arm report4550f8a9890f6efd39f757fcea9c5b7ea870a7beccf9fc84c715d66c7a7f6448

Check it yourself

Download latest.json, the per-model files and the archive, unpack the archive and run these from its folder.

  1. Check the published numbers

    Needs Node 22 or later and nothing else. It recomputes every score, interval, tier and lift from the per-run outcomes, recomputes the hashes from the files in the archive, and re-scores every run on a public case from its stored output.

    Terminal
    node bench/bench.mjs verify --results latest.json --archive paydirt-2026-10-public.tar.gz
    node bench/bench.mjs verify --results practice/latest.json --archive paydirt-2026-10-public.tar.gz
    
  2. Check the public cases

    Needs Foundry for the Solidity proofs. Every planted bug must pass its proof on the vulnerable twin and fail it on the fixed one.

    Terminal
    node bench/bench.mjs lint --proofs
    
  3. Run a model on the practice set

    Needs omp 18.4.4 and an OpenRouter key. Models are not deterministic: a rerun lands near the published numbers, not on them.

    Terminal
    node bench/bench.mjs run --set public --models <slug> --run-id mine
    node bench/bench.mjs score --set public --run-id mine
    

The practice set

6 public pairs, published with their answer keys, proofs and the raw output of every run on them, so every outcome here can be re-scored end to end. This is not the Paydirt Score.

ModelPairs rightRecallFool’s goldCost per run
GPT-6 LunaPairs right6 of 6Recall100%Fool’s gold0%Cost per run$0.0086
DeepSeek V4.1 FlashPairs right5 of 6Recall100%Fool’s gold0%Cost per run$0.02
GLM 5.3 FlashPairs right4 of 6Recall100%Fool’s gold0%Cost per run$0.011

practice/latest.json

Not run

3 models of the release’s list have no score. A model with no answer is listed here, not given a low score.

WhyModels
Its hosts returned tool calls as plain text on 35 of 36 initial inputs. The batch could not be completed and is excluded from the ranking.Model
  • gpt-oss-120bopenai/gpt-oss-120bRun tier 1
OpenRouter serves it only to accounts that have made an age confirmation, which the benchmark account has not made.Model
  • Muse Spark 1.3meta/muse-spark-1.3Run tier 2
Its runs were not finished when this release was published: 35 of 36 inputs completed.Model
  • Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-liteRun tier 2

Release notes

Google Gemini 3.5 Flash Lite has 35 finalized raw inputs and one unresolved input after repeated provider failures. Its results are shown for inspection, but it is not ranked or recommended.

Questions

Why 18 pairs and one run per input?

Each pair is an original case with an executable proof for every planted bug and every decoy, and the scored pairs are held, so the set grows slowly. With 18 pairs the 95% interval is wide; it is published so that a two-pair difference is not read as a ranking. Every model answers every input once. The release’s credit covered one repeat for many models rather than three for a handful, so the score carries the noise of one run per input: the interval covers which pairs were drawn, not how a model varies between runs.

Why is a model missing from the table?

Models run in tiers, the cheapest first inside a tier, until the release’s credit is used. A model the credit did not reach, or whose host returned no answer, is listed under Not run with the reason. A model whose runs had not finished at publication is listed there too, with how many of its inputs were completed. None of them is given a low score.

Is the benchmark tuned to Bounty Operator?

Bounty Operator is not a row. The ranked score comes from the raw arm, which sends no product text: the same system prompt, task and answer sheet to every model. The profile arms, which add a Bounty Operator core profile, are published separately, negative lifts included. Each pair started as a draft by a model of one of nine vendors. Sessions of Claude Opus 5.5, a ranked model, directed by the benchmark’s owner, then checked, repaired and proved every pair and wrote two of the hard pairs. Those two count as Anthropic-drafted, and the results file gives every model’s score without the pairs its own vendor drafted. Two ranked models, DeepSeek V4.1 Flash and GLM 5.3 Flash, ran the pilot on the find pairs of that time, and while nine pairs were written, drafts were sent to a few ranked models to see whether they could be decided both ways; the method names the models and the pairs whose design changed after such a check.

What does held mean?

A held case’s files are not published. Before the first scored run each held case got a salted SHA-256 commitment (36 in this release); when a held case moves to the public set, its salt is published with it so the commitment can be checked. Held does not mean unseen by the model providers: a held case is sent to the providers that serve each model, as every prompt is.

When does it update?

With each release. A release spends a fixed amount of credit in run order: tier 1 first, the cheapest model first inside a tier. Held pairs move to the public set over time and new held pairs replace them. A change to the harness, the routing policy, the task text or a set changes every hash, and no earlier run counts after it.