Settle & Tape
Independent forensics for trading bots

Your bot says it’s profitable. Let’s check the tape.

We rebuild your bot’s PnL from the only sources that can’t flatter it — official market settlement, exchange fills and on-chain cash flow — then try to break every edge it claims before you put more money behind it.

Polymarket · Binance spot & perps · Deribit options · custom Python bots. Read-only exports only — we never ask for keys.

Exhibit A — Futures momentum bot Restated
Total printed by the bot343 closed trades on perpetual futures +$53.66
Positions restored after a 93-day pauseOpened in June, closed at September prices on restart −$60.30
Stats counter reset counted twice −$9.11
Exchange fees the code never charged5 bp per side on every fill −$38.20
Verified net result −$53.95
From our internal fleet audit. Per-trade mean −$0.157, 95% CI [−0.253, −0.061], t = −3.21. Verdict: RETIRE.
Why this exists

We audited our own 18-bot fleet first. The ledgers were wrong by $1,685.97.

Before offering this to anyone, we ran it on ourselves: 6,452 settled trades reconciled against official outcomes, exchange data and the wallet. Almost none of the error was exotic. It was bookkeeping.

$1,685.97absolute gap between bot ledgers and verified truth · $0.16 left unexplained
Duplicate settlement rows91.4%
Stale reference price used to settle7.1%
Fees charged on winning trades only0.8%
Wrong self-settlement & phantom fills0.74%
What we check

Eight things that turn a paper profit into a real loss.

Every audit scores your bot on the same eight dimensions, so you can compare versions, strategies and vendors on one sheet.

C1

Accounting truth

Duplicate rows, fees applied to winners only, the wrong settlement source, positions restored across restarts.

C2

Data integrity & look-ahead

Unclosed candles in features, timestamps off by a bar, survivorship in the symbol universe, dropped days in the backtest.

C3

Execution realism

Fills at mid or last price, maker fills that never happened, ghost fills that matched off-chain but failed on-chain, adverse selection, and no-fills near expiry.

C4

Statistical evidence

Out-of-sample split, day-clustered bootstrap intervals, multiple-testing correction, and how many trades a verdict actually needs.

C5

Live safety & state

Kill switch, risk limits, order idempotency, orphaned positions after a crash, what happens when an order’s fate is unknown.

C6

Real-money reconciliation

Your ledger against exchange history or on-chain cash flow, market by market, until every dollar has a named cause.

C7

Operational health

Stale feeds treated as live, crash loops, silent restarts, heartbeats that report alive while nothing trades.

C8

Claims vs reality

The numbers in your README, pitch deck or vendor sales page, reproduced exactly — then re-run the way the bot really trades.

Case files

What the numbers looked like before and after.

All from our own fleet, anonymized. None of these bots crashed or threw errors. Each one simply reported an edge that wasn’t there.

Options · weekly short straddle

The backtest skipped the violent weeks

Backtest+113 bp
All Fridays+23 bp

A liquidity filter dropped 26 of 104 Fridays — the ones that moved most. Restored at the same Deribit costs, t fell from 1.83 to 0.34 and the worst week doubled to −25.5% of notional.

Perps · session breakout

512 trades that were really 200

Reported t2.65
Clustered by day1.55

335 entries opened on the same bar as another trade. Removing the top 10 trades cut the mean from +57.9 bp to +12.5 bp.

Prediction markets · live wallet

The risk halt that wasn’t a breach

Risk gate−$15.11
On-chain−$11.97

Meanwhile the bots’ own records were $13.62 too optimistic: 16 markets traded but never logged, and 7 logged at the wrong value.

Process

Reproduce first. Then try to break it.

The order matters: we never argue with a number we couldn’t reproduce, and we register what would count as failure before we look at the results.

  1. Intake

    You send code, ledgers and read-only exports. We agree scope and the claims under test in writing.

  2. Reproduce

    We rebuild your headline numbers exactly. If we can’t, that is the first finding.

  3. Rebuild truth

    Every trade re-settled against the official source, exchange fills or on-chain cash flow, with fees applied to every fill.

  4. Attack the edge

    Pre-registered tests: look-ahead, holdout, day-clustered bootstrap, cost stress, a bot-faithful replay.

  5. Review live safety

    Crash, restart and unknown-order scenarios simulated offline against your code — never against your account.

  6. Verdict & fix list

    A scored report, a ranked list of fixes with file and line, and a clear call on what to do with the bot.

The deliverable

One scorecard, one verdict, every claim traced.

DimensionScore
Code correctness4 / 5
Data integrity & look-ahead2 / 5
Accounting truth3 / 5
Execution realism4 / 5
Live safety & state4 / 5
Operational health4 / 5
Statistical evidence1 / 5
Claims vs reality1 / 5
Verdict
KEEPIMPROVERESEARCHPAPERLIVE SMALLRETIRE
Scorecard & verdictEight dimensions, 0–5, with the evidence behind each score.
Restated PnLPer trade, per day, with every adjustment named — like Exhibit A.
FindingsConfirmed vs suspected, each with file and line, and how we proved it.
Fix listRanked by money at risk, with a test that shows each fix worked.
Sample-size planHow many trades or days until the edge can be judged, and a stop rule.
ScriptsThe reconciliation and test scripts, so you can re-run them yourself.
Pricing

Fixed scope, fixed price.

Priced per bot. Larger fleets and firms with multiple strategies get a written quote after intake.

Ledger check
$750per bot

Is the PnL real?

  • Accounting truth (C1) and reconciliation (C6)
  • Restated PnL with named adjustments
  • Top findings with file and line
5 business days · up to 10,000 trades
Full forensic audit
$2,900per bot

Is the edge real, and is it safe to run?

  • All eight dimensions, scored
  • Pre-registered statistical tests and bot-faithful replay
  • Live-safety simulation and ranked fix list
  • Sample-size plan, stop rule and scripts
  • One 60-minute findings call
10–15 business days
Ongoing verification
$1,200per month

Keep it honest after launch.

  • Weekly reconciliation against exchange or chain
  • Stop-rule checks on frozen criteria
  • Re-audit of each change before it goes live
Up to 3 bots · after a full audit

Payment in USDC on Base, Polygon, Ethereum or Solana, or by bank transfer. 50% on signing, 50% on delivery. Invoices state the exact network and amount.

Boundaries

What we will never do

  • Never ask for API keys, seed phrases or private keys.
  • Never place, cancel or modify orders on your account.
  • Never sell signals, manage funds or promise returns.
  • Never tell you a bot is profitable without holdout evidence.

Audits are technical reviews of software and data, not investment, legal or tax advice, and never a recommendation to trade any asset. A clean audit is not a guarantee of future results.

Case-file figures are historical results of paper, demo and small live bots, some simulated. Simulated results have inherent limits: they do not reflect real execution and may under- or over-state the effect of market factors such as liquidity. Trading involves substantial risk of loss.

What to send

Everything read-only

  • Bot source code, or a repository with read access
  • Trade ledgers or logs, as they are — don’t clean them
  • Exchange trade history export, or your public wallet address
  • The backtest and any numbers you’ve published or been sold

Find out what your bot really made.

Tell us the venue, the strategy in one sentence and the number you want verified. We reply within two business days with scope and a fixed quote.

Request an audit Venue · strategy · trades to date · claimed result