Beyond one backtest¶
One backtest is one path through one tape. It cannot tell you whether the edge is real, whether it survives forward, or whether the pass you just saw was luck. Four tools address four different failure modes, and none of them substitutes for another.
| Tool | The question it answers | What it cannot rescue |
|---|---|---|
| Sequential combines | How did this do across real, separate attempts? | A short tape — you get few attempts |
| Monte Carlo | Given these daily outcomes, how often does a combine pass? | A regime the tape never contained |
| MC confidence | How wrong could that probability be? | The uncalibrated rule constants |
| Walk-forward | Does the edge survive forward in time? | A selection rule that never generalized |
| PBO | Does my selection generalize at all? | Decay — it is symmetric in time |
| Deflated Sharpe | Is this Sharpe real after the trials I ran? | Trials you did not record |
Sequential combines¶
Nobody runs a single continuous evaluation for three years. They attempt a combine, and if it
resolves, they attempt another. metrics.sequential_combines models exactly that: it cuts a
long tape into consecutive non-overlapping windows of N trading days and runs each as a
completely fresh evaluation — new balance, new floor, new broker, new strategy instance.
from topstep_backtest.metrics import sequential_combines
sweep = sequential_combines(bars, MyStrategy, window_days=21) # ~30 calendar days
print(sweep.attempts, sweep.pass_rate, sweep.mll_breach_rate)
for w in sweep.windows:
print(w.index, w.first_day, w.last_day, w.verdict.name, w.failure, w.net_pnl)
factory is the strategy class or a zero-argument callable, never an instance: engine,
broker, kernel and strategy state are all single-use, so a reused instance would carry indicator
state across windows and destroy the independence the whole exercise depends on.
Non-overlapping is the point, not a limitation. Stepping the window one day at a time would give roughly 700 windows from three years, each sharing 20 of its 21 days with its neighbour. That pass rate would carry the precision of 700 observations and the information of about 35. The step is fixed at the window length rather than shipped as a footgun with a warning attached.
Read pass_rate beside attempts. Three years is about 35 attempts, so the standard error
on a pass rate is near eight percentage points before anything else is considered. It is an
estimate off a handful of samples, not a probability.
Warm starts, and why they happen outside the engine¶
Each window gets a fresh strategy, so its indicators would normally start cold and gate away the
window's opening bars — every window measured partly during its own transient. So the history
immediately preceding a window is driven through SymbolStrategy.prewarm before the run.
That has to happen outside the engine. Feeding preload bars through the Backtest instead would
land those days in day_records as flat days, which moves closed_days, every daily
percentile, stdev, sortino and the drawdown durations while leaving P&L untouched — a
report that is wrong in exactly the places nobody checks.
The preload size is history_bars, not warmup — where the value stops depending on where the
run started, rather than where it merely exists. That is a real cost: Sma(30) is warm at 1,920
bars, which on RTH-only 5-minute candles is about 25 trading days of history per window.
The opening windows of any tape have no such history and are reported fully_warm=False rather
than quietly counted as equals. Cold windows under-trade, so leaving them in biases the pass
rate down.
trailing_days_dropped reports the days left over by the last whole window. A partial window is
not an attempt, and scoring one would count a short evaluation as a failure to reach the target.
Run it beside the Monte Carlo¶
These two estimate the same quantity by opposite methods and are biased in opposite directions. Monte Carlo resamples the observed days — many paths, but real sequencing destroyed beyond the block length; it cannot know your worst three days were consecutive because they were one news cycle. Sequential combines preserve sequencing and regime exactly and pay for it in sample size.
They share classify_failure, so the autopsies are directly comparable. Run both — and
metrics.confidence.crosscheck(mc, sweep) does the comparison for you, outcome by outcome,
flagging any rate the real windows put outside 2×SE of what the bootstrap's own probability
would produce over that many attempts. When they disagree, the disagreement is the
finding — a pass rate that is much better in the bootstrap than in the real windows
usually means your losses cluster in a way the block length missed.
Monte Carlo¶
metrics.monte_carlo block-bootstraps the run's own observed trading days and replays each
synthetic sequence through a fresh CombineKernel — the same rule engine the live backtest
used, not a reimplementation of it.
from topstep_backtest.metrics import monte_carlo
mc = monte_carlo(report.result, params=report.params, paths=10_000, seed=7)
Read the autopsy, not the pass probability. The headline number alone tells you nothing actionable. The failure partition does:
mll_breach— you are too big. Resize.consistency_blocked— one day is carrying too much of the profit. Throttle it.target_not_reached— the edge is too slow for the window. Neither of the above will help.
Those three imply completely different fixes, and a single probability hides which one you have.
block_length=1 is a footgun
It degenerates to an i.i.d. resample, destroys the losing streaks that actually blow accounts, and reports a pass probability that is far too kind. The default is 5.
Two more constraints worth internalizing. The horizon defaults to one billing month
(BILLING_MONTH_DAYS = 21) — a combine has no time limit, only a monthly fee, so "one
attempt" means one fee cycle, and the fixed default is never the observed day count: a
blown run stops recording days at the breach, that count is a survival time rather than a
combine length, and source_truncated flags such a sample as survivorship-biased by
construction. And it is provisional below 30 source days: resampling cannot create
information that was not in the sample.
How much to trust that number¶
With thousands of paths the simulation error is negligible — which makes the number easy
to over-read. The real uncertainty is the day sample, the block assumption, and the regimes
the tape never held. metrics.confidence quantifies each, and every estimate runs through
the same core as the number it qualifies:
from topstep_backtest.metrics.confidence import mc_confidence
c = mc_confidence(report.result, params=report.params, seed=7)
c.ci.p05, c.ci.p95 # double-bootstrap band: the error bar the DAY COUNT earns
c.sensitivity.spread # wide = the streak assumption is doing the work
c.by_year.spread # the error bar non-stationarity imposes
Quote the CI, not the point. A 65% from 500 source days and a 65% from 40 are different
findings, and the band is what says so. The per-year strata refuse to average a hostile
year against a kind one, and the block-length row shows whether the one arbitrary knob was
load-bearing. The whole bundle renders on the tearsheet:
report.to_html(path, confidence=c, crosscheck=check). None of it repairs the epistemic
caveats — the constants stay uncalibrated and no resample invents a missing regime; these
bound the statistical error only.
Optimization, and why it needs guards¶
optimize() was deliberately absent from this project for a long time. Maximizing over a
combine metric is an overfitting machine. It exists now because the guards exist.
from topstep_backtest.metrics import optimize, walk_forward
from topstep_backtest.metrics import deflated_sharpe, probability_of_backtest_overfitting
OptimizationResult.best is not evidence. Selecting a maximum from a search guarantees a
flattering number. The API keeps every trial precisely so you can feed pbo_matrix to
probability_of_backtest_overfitting and the winner's daily_pnl to deflated_sharpe. A
bare best-params answer is the thing to refuse — including from yourself.
The default objective is net P&L, deliberately not the pass/fail verdict: optimizing a binary throws away nearly all the information in a run and rewards configurations that scraped over the line once.
Cost is honest and serial by design — a grid of G over F folds runs G × (F + 1)
complete backtests. A search that silently sampled would be worse than a slow one.
Walk-forward¶
Anchored walk-forward asks whether the edge survives forward.
Read efficiency, not the out-of-sample number. The question is how much of the
in-sample edge survived. Negative efficiency — chosen for making money, then lost money — is
the signature of a fit to noise. It is None against a non-positive in-sample objective,
because a ratio against a loss inverts the sign of good and bad.
consistent_folds guards the aggregate: a good total built from one huge fold and three
losers is not a robust strategy, and the total alone will not tell you.
PBO¶
Probability of Backtest Overfitting, by combinatorially symmetric cross-validation. It asks whether your selection procedure generalizes — whether the configuration that won in-sample tends to be below median out-of-sample.
PBO and walk-forward are not substitutes. PBO is symmetric in time, so it says nothing about decay. Walk-forward is directional, so it conflates a decaying edge with a bad selection rule. Run both. They disagree in informative ways, and the disagreement is the finding.
Deflated Sharpe¶
A Sharpe ratio computed after twenty trials is not the same evidence as one computed after a
single hypothesis. deflated_sharpe corrects for the number of attempts, read from a
TrialLedger you keep as you search.
The ledger only knows what you record. Trials you ran and forgot — including the ones you abandoned because they looked bad — are exactly the ones that make the correction matter, and no amount of statistics can recover them after the fact.
Economics: the question that actually decides it¶
A pass probability is not a decision. metrics.evaluate_ev turns one plus your own prices
into an expected value per attempt, and a breakeven pass rate:
from topstep_backtest.metrics import evaluate_ev
ev = evaluate_ev(pass_probability=mc.pass_probability, economics=my_economics)
breakeven_pass_rate is the number to reason with: the pass probability at which the whole
exercise turns EV-neutral. Comparing it against the Monte Carlo estimate tells you how much
of your margin is real and how much is inside the estimator's error bars.
Multi-year runs¶
A two-year run spans roughly eight quarterly rolls, and it needs
stitch_continuous.
The counter-intuitive part is worth stating plainly: a raw splice does not produce wrong
P&L. The 16:10 flatten plus the session-roll backstop mean no position and no working order
survives a day boundary, and a roll seam is a day boundary — so every entry and its exit
share one contract. What the seam corrupts is indicator state, which does span it: a
spurious Cross, an Atr spike inflated for a whole lookback, a false breakout. Those
trades are priced correctly and should never have been taken.
Additive back-adjustment is therefore exactly P&L-neutral, because entry and exit share the
offset and it cancels in the difference. Never ratio-adjust: multiplicative offsets do
not cancel, and they push prices off the tick grid that SimBroker asserts on at every fill.
What adjustment does distort is logic keyed to absolute price levels — round numbers, a
fixed price target. Tick-relative logic (stop_loss_ticks, take_profit_ticks, all the
strategy sugar) is unaffected.
What none of this fixes¶
Every tool on this page inherits the operating limits
of the backtest that fed it. Monte Carlo cannot invent a regime the tape never contained.
PBO cannot correct a look-ahead you introduced with the wrong stamp. Walk-forward on
holiday-contaminated data walks forward through sessions that never existed.
And all of them inherit the uncalibrated rule and fee constants. A pass probability quoted to three decimals from unverified inputs is precise, not accurate.