Operator

One run, eight stages, one verdict on your bug bounty report

Hand over the draft, the code and the programme rules. The Gauntlet runs eight stages in the order that saves work and returns one decision, the one blocker behind it and the cheapest action that removes it.

US$10 per week. Runs on your own model and key.

ExampleExample dossier · invented protocol

VerdictProve first

The seizure bug is in the code. The proof reads the return value of liquidate on a mock oracle; the impact row names the borrower’s collateral.

  • 0Critical
  • 1High
  • 0Medium
  • 1Hardening
  • 3Checked safe
Blocker
The proof never reads the borrower’s collateral balance, and the unhealthy state comes from a mock oracle.
Cheapest action
One fork test at the deployed revision: let the real feed move the price, liquidate with 1 unit of debt, assert the borrower’s collateral before and after, and run a control that repays in full.
Severity to claim
High, on the programme scale supplied in Context.
Filing deadline
Within hours of the fork run. The fix is one line on the central liquidation path.

The verdict is one of five

Words a hunter already uses. The dossier never ends on “consider” or “it depends”.

  • One verdict for the whole report.
  • One blocker: the single fact that stops submission today.
  • The cheapest action that removes the blocker, as something you run locally.
  • A filing deadline, set by how exposed the finding is to a private duplicate.
Submit
Every decisive claim is backed by the supplied code and proof. File it.
Rewrite, then submit
The finding holds. The draft misstates its severity, impact, preconditions or fix.
Prove first
The claimed impact is not demonstrated on the real path yet. One artefact is missing.
Hold: duplicate
A known issue, an audit note, a branch or your own earlier report has the same root cause or the same fix.
Drop
The code contradicts the root cause, the behaviour is the design, or the rules exclude it.

Eight stages, gates first

Scope, provenance and prior art end a report without a proof. They run before the stages that cost you work. Each stage reads the output of the stages before it.

Stage 1

Scope

Whether the code is in the scoped asset at the deployed revision, whether the class is excluded, and whether every clause of the chosen impact row has an artefact behind it.

  • Binding table
  • Exclusion matches
  • Impact row, clause by clause

Stage 2

Provenance

Who performs each step, whether the behaviour is documented design or an audit fix, whether the bug adds anything over the intended path, and whether every precondition is reachable from live state.

  • Actor table
  • Intent evidence
  • Counterfactual
  • Preconditions

Stage 3

Prior art

Whether a known issue, a prior audit note, a team branch or your own earlier report has the same root cause or the same one-line fix.

  • Root-cause fingerprint
  • Matches
  • The duplicate clock

Stage 4

Proof

Whether the proof runs production code on every step and ends by reading the object the impact row names, with measured numbers.

  • Step table: executed, mocked or narrated
  • End-state assertion
  • Measured loss

Stage 5

Severity

The highest row of the programme’s own scale that the proof fully asserts, after every downgrade clause.

  • The tier to claim
  • The one fact that moves it

Stage 6

Triager

The one sentence that closes this report in ten minutes, and whether the first paragraph already answers it.

  • Ranked rejection reasons
  • The draft sentence that triggers each

Stage 7

Report

Whether the claims follow from the code, the proof is inline, the form and the body agree, and the limits are stated.

  • Claims table
  • The rewritten report

Stage 8

Verdict

One decision.

  • One of five verdicts
  • One blocker
  • The cheapest action that removes it
  • A filing deadline

The dossier, in full

A draft claims a Critical pool drain in Brinewell Lend, a lending protocol invented for this page. Eight stages later the finding stands, the severity moves, and one test is missing.

ExampleExample dossier · invented protocol · Brinewell Lend

VerdictProve first

The seizure bug is in the code. The proof reads the return value of liquidate on a mock oracle; the impact row names the borrower’s collateral.

  • 0Critical
  • 1High
  • 0Medium
  • 1Hardening
  • 3Checked safe
Stages
StageIts verdictThe one thing it decided
1ScopeIts verdictSubmitThe one thing it decidedThe file and the function exist in the scoped asset at the deployed revision. No exclusion names the class.
2ProvenanceIts verdictProve firstThe one thing it decidedEvery step is performed by an unprivileged liquidator. The unhealthy state is set through MockOracle.setPrice, with no route from live state shown.
3Prior artIts verdictRewrite, then submitThe one thing it decidedAudit item L-04 has the same symptom and a different root cause. The report has to state the difference in its first paragraph.
4ProofIts verdictProve firstThe one thing it decidedThe final assertion reads the return value of liquidate. The impact row names the borrower’s collateral.
5SeverityIts verdictRewrite, then submitThe one thing it decidedThe proof supports High on the scale supplied. The draft selects Critical.
6TriagerIts verdictProve firstThe one thing it decidedFastest close: impact not shown. “The attacker drains the pool” has no assertion behind it.
7ReportIts verdictRewrite, then submitThe one thing it decidedRoot cause and fix are confirmed against lines 215-216. The title claims a pool drain; the code shows a loss per position.
8VerdictIts verdictProve firstThe one thing it decidedOne fork test stands between this draft and a High.
input-2/src/LiquidationEngine.solsha2567c1e9a40b6d2…318 lines
F-1HighProven in source

liquidate sizes the seizure from the whole collateral, not from the amount repaid

  • input-2/src/LiquidationEngine.sol:212-221
Impact

A liquidator who repays 1 unit of debt on an unhealthy position takes all of its collateral. The borrower loses the collateral above the debt. Bound: each unhealthy position’s collateral, minus the debt repaid.

Observed
  1. liquidate checks one thing about the position: _healthFactor(p) < WAD (line 214).
  2. Line 215 computes seized from p.collateral and bonusBps. repay is not in the expression.
  3. Line 216 caps seized at p.collateral, so any repay above zero seizes the whole collateral.
  4. Lines 217-220 reduce the debt by repay and send seized to the caller.
input-2/src/LiquidationEngine.sol212–221
function liquidate(address borrower, uint256 repay) external nonReentrant {
    Position storage p = positions[borrower];
    if (_healthFactor(p) >= WAD) revert Healthy();
    uint256 seized = (p.collateral * (BPS + bonusBps)) / BPS;
    if (seized > p.collateral) seized = p.collateral;
    p.debt -= repay;
    p.collateral -= seized;
    debtToken.safeTransferFrom(msg.sender, address(this), repay);
    collateralToken.safeTransfer(msg.sender, seized);
}
Counterargument

The position is unhealthy. The borrower was going to lose that collateral to liquidation anyway.

OpenThe intended path seizes the repaid amount plus the bonus. The draft never shows the two outcomes side by side, and its test reaches the unhealthy state through MockOracle.setPrice.

Evidence gap
  • A run on a fork of the deployed revision where the position turns unhealthy through the real price feed.
  • A final assertion on the borrower’s collateral balance, next to a control run that repays the debt in full.
Fix

Compute seized from repay: convert it to collateral at the oracle price, add bonusBps, then cap at p.collateral.

Next

Run the fork test with both assertions, then paste the command and its output into the report body.

To do, in order
#ActionArtefact it producesStage
1ActionRerun the test on a fork of the deployed revision, with the price moved by the real feed.Artefact it producesCommand and outputStageProof
2ActionAssert the borrower’s collateral balance before and after, and add the full-repay control.Artefact it producesTwo assertionsStageProof
3ActionChange the claimed severity to High and quote the impact row word for word.Artefact it producesEdited reportStageSeverity
4ActionName audit item L-04 in the first paragraph and state the different root cause.Artefact it producesOne paragraphStagePrior art
5ActionReplace “drains the pool” in the title with the per-position loss.Artefact it producesEdited titleStageReport

What you hand over

The run asks for the evidence that decides outcomes. A field left empty is named in the dossier as not supplied.

  1. The draft and the code

    Your report, the source files it cites and the proof with its command and output.

  2. Impact and exclusions

    The programme’s impact list, the row you intend to select, the exclusions and the trusted roles.

  3. The severity scale

    The programme’s own scale with its thresholds and downgrade clauses, pasted in.

  4. Asset and revisions

    The scoped asset, the revision your proof ran against and the revision that is deployed.

  5. Prior material

    Known issues, audits and fix-review notes, team branches, and your own earlier reports on the programme.

  6. Clock and read-back

    The date you first reproduced it, the fee and duplicate rules, and what the platform stored after you filled in the form.

How it runs

No new machinery. The Gauntlet is eight hosted reviews, run in order from your browser on your own key.

  1. Eight reviews, one after another

    Each stage is a hosted review with its own profile. It goes through our server, which adds the method of the stage, to your provider. Your provider bills each call to your key.

  2. Earlier stages become evidence

    A stage’s output goes to the later stages as a file named stage-<n>-<profile>.md, next to your draft and your code.

  3. Finished stages are kept

    A cancelled run, a provider error or a reload keeps every stage that finished. The run picks up at the stage that did not.

  4. One packet

    The dossier downloads as one record with the SHA-256 manifest of every file the run read.

Questions

What does a gauntlet run cost?

The Gauntlet is part of Operator at US$10 per week. A run is eight reviews on your own model, so your provider bills eight model calls to your key.

Which model runs the stages?

The one you choose. Every stage is a hosted review: your files and your API key go through our server, which adds the method of the stage, to your provider. That is OpenRouter, Anthropic, OpenAI, Google Gemini, xAI, DeepSeek, Mistral or Groq. No file, key or review is stored.

Can I run the stages on a chat subscription or from a coding agent?

The stages run hosted, on an API key. The report stage is the Challenge a draft report profile, which also exports as a prompt on its own. From a coding agent, the gauntlet prompt of the MCP server runs the same order: seven hosted reviews through run_review with a connection token, and the report stage on your agent’s own model.

What happens when an early stage ends the report?

The run stops there. A drop or a hold-duplicate from scope, provenance or prior art ends the report before the proof stage costs you a day, and the dossier opens on that verdict. One button runs the remaining stages anyway.

Does the gauntlet run code or touch the target?

It reads the files you supply. The proof stage reads your test and its output and tells you the one run that is missing. You run it locally.

What does the free plan show?

The example dossier on this page and every single profile, one hosted review per UTC day. Each gauntlet stage except the final verdict is also a profile you run on its own.

Run the gauntlet before the triager does

Unlimited hosted reviews, the Gauntlet, Panel review and four reviews running at once.

US$10 billed weekly, renews until you cancel in the Stripe portal; access runs to the end of the paid week.