Skip to content

Beyond one backtest

One backtest is one path through one tape. It cannot tell you whether the edge is real, whether it survives forward, or whether the pass you just saw was luck. Four tools address four different failure modes, and none of them substitutes for another.

Tool The question it answers What it cannot rescue
Sequential combines How did this do across real, separate attempts? A short tape — you get few attempts
Monte Carlo Given these daily outcomes, how often does a combine pass? A regime the tape never contained
MC confidence How wrong could that probability be? The uncalibrated rule constants
Walk-forward Does the edge survive forward in time? A selection rule that never generalized
PBO Does my selection generalize at all? Decay — it is symmetric in time
Deflated Sharpe Is this Sharpe real after the trials I ran? Trials you did not record

Sequential combines

Nobody runs a single continuous evaluation for three years. They attempt a combine, and if it resolves, they attempt another. metrics.sequential_combines models exactly that: it cuts a long tape into consecutive non-overlapping windows of N trading days and runs each as a completely fresh evaluation — new balance, new floor, new broker, new strategy instance.

from topstep_backtest.metrics import sequential_combines

sweep = sequential_combines(bars, MyStrategy, window_days=21)  # ~30 calendar days
print(sweep.attempts, sweep.pass_rate, sweep.mll_breach_rate)
for w in sweep.windows:
    print(w.index, w.first_day, w.last_day, w.verdict.name, w.failure, w.net_pnl)

factory is the strategy class or a zero-argument callable, never an instance: engine, broker, kernel and strategy state are all single-use, so a reused instance would carry indicator state across windows and destroy the independence the whole exercise depends on.

Non-overlapping is the point, not a limitation. Stepping the window one day at a time would give roughly 700 windows from three years, each sharing 20 of its 21 days with its neighbour. That pass rate would carry the precision of 700 observations and the information of about 35. The step is fixed at the window length rather than shipped as a footgun with a warning attached.

Read pass_rate beside attempts. Three years is about 35 attempts, so the standard error on a pass rate is near eight percentage points before anything else is considered. It is an estimate off a handful of samples, not a probability.

Warm starts, and why they happen outside the engine

Each window gets a fresh strategy, so its indicators would normally start cold and gate away the window's opening bars — every window measured partly during its own transient. So the history immediately preceding a window is driven through SymbolStrategy.prewarm before the run.

That has to happen outside the engine. Feeding preload bars through the Backtest instead would land those days in day_records as flat days, which moves closed_days, every daily percentile, stdev, sortino and the drawdown durations while leaving P&L untouched — a report that is wrong in exactly the places nobody checks.

The preload size is history_bars, not warmup — where the value stops depending on where the run started, rather than where it merely exists. That is a real cost: Sma(30) is warm at 1,920 bars, which on RTH-only 5-minute candles is about 25 trading days of history per window. The opening windows of any tape have no such history and are reported fully_warm=False rather than quietly counted as equals. Cold windows under-trade, so leaving them in biases the pass rate down.

trailing_days_dropped reports the days left over by the last whole window. A partial window is not an attempt, and scoring one would count a short evaluation as a failure to reach the target.

Run it beside the Monte Carlo

These two estimate the same quantity by opposite methods and are biased in opposite directions. Monte Carlo resamples the observed days — many paths, but real sequencing destroyed beyond the block length; it cannot know your worst three days were consecutive because they were one news cycle. Sequential combines preserve sequencing and regime exactly and pay for it in sample size.

They share classify_failure, so the autopsies are directly comparable. Run both — and metrics.confidence.crosscheck(mc, sweep) does the comparison for you, outcome by outcome, flagging any rate the real windows put outside 2×SE of what the bootstrap's own probability would produce over that many attempts. When they disagree, the disagreement is the finding — a pass rate that is much better in the bootstrap than in the real windows usually means your losses cluster in a way the block length missed.


Monte Carlo

metrics.monte_carlo block-bootstraps the run's own observed trading days and replays each synthetic sequence through a fresh CombineKernel — the same rule engine the live backtest used, not a reimplementation of it.

from topstep_backtest.metrics import monte_carlo

mc = monte_carlo(report.result, params=report.params, paths=10_000, seed=7)

Read the autopsy, not the pass probability. The headline number alone tells you nothing actionable. The failure partition does:

  • mll_breach — you are too big. Resize.
  • consistency_blocked — one day is carrying too much of the profit. Throttle it.
  • target_not_reached — the edge is too slow for the window. Neither of the above will help.

Those three imply completely different fixes, and a single probability hides which one you have.

block_length=1 is a footgun

It degenerates to an i.i.d. resample, destroys the losing streaks that actually blow accounts, and reports a pass probability that is far too kind. The default is 5.

Two more constraints worth internalizing. The horizon defaults to one billing month (BILLING_MONTH_DAYS = 21) — a combine has no time limit, only a monthly fee, so "one attempt" means one fee cycle, and the fixed default is never the observed day count: a blown run stops recording days at the breach, that count is a survival time rather than a combine length, and source_truncated flags such a sample as survivorship-biased by construction. And it is provisional below 30 source days: resampling cannot create information that was not in the sample.

How much to trust that number

With thousands of paths the simulation error is negligible — which makes the number easy to over-read. The real uncertainty is the day sample, the block assumption, and the regimes the tape never held. metrics.confidence quantifies each, and every estimate runs through the same core as the number it qualifies:

from topstep_backtest.metrics.confidence import mc_confidence

c = mc_confidence(report.result, params=report.params, seed=7)
c.ci.p05, c.ci.p95  # double-bootstrap band: the error bar the DAY COUNT earns
c.sensitivity.spread  # wide = the streak assumption is doing the work
c.by_year.spread  # the error bar non-stationarity imposes

Quote the CI, not the point. A 65% from 500 source days and a 65% from 40 are different findings, and the band is what says so. The per-year strata refuse to average a hostile year against a kind one, and the block-length row shows whether the one arbitrary knob was load-bearing. The whole bundle renders on the tearsheet: report.to_html(path, confidence=c, crosscheck=check). None of it repairs the epistemic caveats — the constants stay uncalibrated and no resample invents a missing regime; these bound the statistical error only.

Optimization, and why it needs guards

optimize() was deliberately absent from this project for a long time. Maximizing over a combine metric is an overfitting machine. It exists now because the guards exist.

from topstep_backtest.metrics import optimize, walk_forward
from topstep_backtest.metrics import deflated_sharpe, probability_of_backtest_overfitting

OptimizationResult.best is not evidence. Selecting a maximum from a search guarantees a flattering number. The API keeps every trial precisely so you can feed pbo_matrix to probability_of_backtest_overfitting and the winner's daily_pnl to deflated_sharpe. A bare best-params answer is the thing to refuse — including from yourself.

The default objective is net P&L, deliberately not the pass/fail verdict: optimizing a binary throws away nearly all the information in a run and rewards configurations that scraped over the line once.

Cost is honest and serial by design — a grid of G over F folds runs G × (F + 1) complete backtests. A search that silently sampled would be worse than a slow one.

Walk-forward

Anchored walk-forward asks whether the edge survives forward.

Read efficiency, not the out-of-sample number. The question is how much of the in-sample edge survived. Negative efficiency — chosen for making money, then lost money — is the signature of a fit to noise. It is None against a non-positive in-sample objective, because a ratio against a loss inverts the sign of good and bad.

consistent_folds guards the aggregate: a good total built from one huge fold and three losers is not a robust strategy, and the total alone will not tell you.

PBO

Probability of Backtest Overfitting, by combinatorially symmetric cross-validation. It asks whether your selection procedure generalizes — whether the configuration that won in-sample tends to be below median out-of-sample.

PBO and walk-forward are not substitutes. PBO is symmetric in time, so it says nothing about decay. Walk-forward is directional, so it conflates a decaying edge with a bad selection rule. Run both. They disagree in informative ways, and the disagreement is the finding.

Deflated Sharpe

A Sharpe ratio computed after twenty trials is not the same evidence as one computed after a single hypothesis. deflated_sharpe corrects for the number of attempts, read from a TrialLedger you keep as you search.

The ledger only knows what you record. Trials you ran and forgot — including the ones you abandoned because they looked bad — are exactly the ones that make the correction matter, and no amount of statistics can recover them after the fact.

Economics: the question that actually decides it

A pass probability is not a decision. metrics.evaluate_ev turns one plus your own prices into an expected value per attempt, and a breakeven pass rate:

from topstep_backtest.metrics import evaluate_ev

ev = evaluate_ev(pass_probability=mc.pass_probability, economics=my_economics)

breakeven_pass_rate is the number to reason with: the pass probability at which the whole exercise turns EV-neutral. Comparing it against the Monte Carlo estimate tells you how much of your margin is real and how much is inside the estimator's error bars.

Multi-year runs

A two-year run spans roughly eight quarterly rolls, and it needs stitch_continuous.

The counter-intuitive part is worth stating plainly: a raw splice does not produce wrong P&L. The 16:10 flatten plus the session-roll backstop mean no position and no working order survives a day boundary, and a roll seam is a day boundary — so every entry and its exit share one contract. What the seam corrupts is indicator state, which does span it: a spurious Cross, an Atr spike inflated for a whole lookback, a false breakout. Those trades are priced correctly and should never have been taken.

Additive back-adjustment is therefore exactly P&L-neutral, because entry and exit share the offset and it cancels in the difference. Never ratio-adjust: multiplicative offsets do not cancel, and they push prices off the tick grid that SimBroker asserts on at every fill.

What adjustment does distort is logic keyed to absolute price levels — round numbers, a fixed price target. Tick-relative logic (stop_loss_ticks, take_profit_ticks, all the strategy sugar) is unaffected.


What none of this fixes

Every tool on this page inherits the operating limits of the backtest that fed it. Monte Carlo cannot invent a regime the tape never contained. PBO cannot correct a look-ahead you introduced with the wrong stamp. Walk-forward on holiday-contaminated data walks forward through sessions that never existed.

And all of them inherit the uncalibrated rule and fee constants. A pass probability quoted to three decimals from unverified inputs is precise, not accurate.