The workflow¶
A backtest number is cheap to produce and expensive to trust. This framework's tools are not a menu — they are a sequence, and each stage exists because the stage before it produces a number that looks finished but is not. The order below is the order in which questions become answerable: run mechanics before verdicts, distributions before probabilities, guards before tuning, prices before decisions.
Each stage ends with a gate — the question you must be able to answer before the next stage's output means anything. The gates are where the discipline lives; the code is short.
| # | Stage | Tool | Gate |
|---|---|---|---|
| 0 | Know the envelope | operating limits | Is this strategy even measurable here? |
| 1 | Data in, declared | Backtest.from_dataframe |
Validator clean, stamp known from vendor docs |
| 2 | Write the strategy | SymbolStrategy |
Brackets at entry, TA-Lib only, warmup understood |
| 3 | One recorded run | record=True + tearsheet |
The strategy did what you meant — verdict ignored |
| 4 | Real attempts | sequential_combines |
Enough attempts to say anything; failure mix known |
| 5 | The distribution | mc_confidence + crosscheck |
CI narrow enough to act on; no unexplained divergence |
| 6 | Tuning (only now) | optimize + guards |
DSR survives the honest trial count; PBO low |
| 7 | Price the attempt | evaluate_ev |
Breakeven pass value is comfortable, not heroic |
Stage 0 — Know the envelope before writing code¶
Some strategies cannot be measured by this engine, and the failure is silent: the run completes, the report is plausible, the number is wrong. Check yours against the operating limits first:
- One contract per run. Feed a single contract's bars (already-continuous data is fine;
stitch_continuousexists if yours is not). - The edge must clear the fill model's pessimism. Tier-0 fills are deliberately harsh — next-bar-open market fills, trade-through limits, stop slippage. A per-trade edge under ~8–10 ticks is dominated by that pessimism, not measured by it. Tick-scalping and limit-heavy passive entries are out of scope.
- Holidays and half-sessions are YOUR problem. No calendar ships, on purpose; filter them upstream or they are scored as ordinary days.
- Every rule and fee constant is cited config, not calibration (§9). Every verdict downstream of here is a diagnostic.
Gate: the strategy's edge, instruments and session behavior fit inside the envelope. If not, no downstream number — however good — is about your strategy.
Stage 1 — Data in, declared and validated¶
bt = Backtest.from_dataframe(
df,
MyStrategy(CONTRACT),
contract_id=CONTRACT,
stamp="open", # from the VENDOR'S docs, never from a trial run
unit=AggregateBarUnit.MINUTE,
unit_number=5,
)
stamp, unit and unit_number have no defaults because guessing the stamp is a silent
one-bar look-ahead that inflates every result, and a wrong stamp on RTH-only data is
byte-identically silent. Validation is strict by default — let it refuse. And check the
contract id twice: a wrong --symbol validates cleanly and is ~25× wrong (wrong tick
value, no error anywhere).
Gate: the validator passes, and you know the stamp from documentation rather than from which setting produced nicer results.
Stage 2 — Write the strategy against the dialect¶
Subclass SymbolStrategy (quickstart,
tutorial). Three habits pay for themselves downstream:
- Register every indicator with
use()— TA-Lib only, updated once per bar, warmup gated. No hand-written formulas; none can drift from the reference implementation. - Attach the bracket at entry (
stop_loss_ticks=/take_profit_ticks=). A stop at entry is what defines the trade's initial risk, and initial risk is what makes every R-multiple in the report a true R rather than a retrofitted one. - Narrate with
self.note(). Notes are pure observation — they cannot change a result — and they are what makes stage 3's replay legible: the tape shows what happened, the notes show what you thought was happening.
Gate: you can state the strategy's intent per bar without reference to hidden state, and
you know its warmup cost (history_bars — real money at ~25 trading days of 5-minute RTH
history for a 30-bar indicator).
Stage 3 — One recorded run, read for mechanics, not verdict¶
report = Backtest(bars, MyStrategy(CONTRACT), record=True).run()
print(report)
report.to_html("run.html") # results tab + the replay · bar by bar tab
The single run's job is not to tell you whether the strategy passes. It is one sample — acting on it is how a strategy that got lucky once becomes a funded account that blows up. Its job is to prove the mechanics:
REJECTEDcount is zero or every rejection is explained. A strategy whose orders were refused reports a clean-looking run that measured nothing.- Warmup gating matches expectation —
bars_gatedshould equal your indicator math, not surprise you. - The replay tab shows intent matching execution: step to each entry, check the bracket is where you meant it, watch the notes against the fills.
- The fills are believable against how each metric is based.
Gate: the strategy demonstrably did what you meant. Only now is a verdict worth computing — on more than one sample.
Stage 4 — Stop trusting one sample: real attempts¶
from topstep_backtest.metrics import sequential_combines
sweep = sequential_combines(bars, lambda: MyStrategy(CONTRACT), window_days=21)
Sequential combines cut the tape into
consecutive fresh evaluations — real sequencing, real regimes, prewarmed indicators. Expect
the head of the tape fully_warm=False, and expect TARGET_NOT_REACHED to dominate: that
is the ordinary outcome of a 21-day window, not a failure signal. Read pass_rate beside
attempts — three years is ~35 windows, a standard error near 8 points before anything
else is considered.
Gate: you have enough attempts for the rate to mean anything, and you know which way the failing windows failed.
Stage 5 — The distribution, quoted with its confidence¶
from topstep_backtest.metrics import crosscheck, mc_confidence
c = mc_confidence(report.result, params=report.params, seed=7)
check = crosscheck(c.mc, sweep) # horizons must match: both one billing month
report.to_html("run.html", confidence=c, crosscheck=check)
The Monte Carlo turns your observed days into thousands of synthetic attempts through the real rule kernel; the confidence instruments say how much the resulting number is worth. Read them in this order:
- The autopsy before the probability —
mll_breach/consistency_blocked/target_not_reachedimply different fixes (resize / throttle the outsized day / better edge), and the fix loops you back to stage 2, not to a bigger position. - The CI before the point. A 65% from 500 source days and a 65% from 40 are different findings. Quote the band.
- The block-length spread — wide means the streak assumption is doing the work, and the honest report is the range.
- The year strata — the pooled number averages a hostile year against a kind one; the spread across years is the error bar non-stationarity imposes.
- The crosscheck — the bootstrap and the real windows estimate the same quantity with opposite biases. When they disagree, the disagreement is the finding.
Gate: the CI is narrow enough to act on, nothing is flagged provisional, and the
crosscheck shows no unexplained divergence. If the CI spans "great" to "hopeless", the
answer is a longer tape, not a decision.
Stage 6 — Tuning, only now, and only with the guards on¶
If — and only if — the mechanics are proven and the distribution says the edge is real but mis-sized or mis-tuned:
from topstep_backtest.metrics import optimize, walk_forward
from topstep_backtest.metrics import (
TrialLedger,
deflated_sharpe,
probability_of_backtest_overfitting,
)
optimize keeps every trial
precisely so the guards can price the search: feed the trial matrix to PBO, the winner's
daily P&L and the honest trial count (the TrialLedger, which remembers the sweeps you
abandoned) to deflated Sharpe, and confirm with anchored walk_forward that the edge
survives forward. Then re-run stages 4–5 on the tuned configuration — a tuned strategy
is a new strategy.
Gate: DSR survives at the recorded trial count and PBO is low. A best-params answer without those two numbers is the thing to refuse — including from yourself.
Stage 7 — Price the attempt¶
from topstep_backtest.metrics import EvalEconomics, evaluate_ev
ev = evaluate_ev(
c.mc,
economics=EvalEconomics(
monthly_fee=Decimal("165"), # YOUR prices — nothing is baked in
pass_value=Decimal("2500"), # your honest estimate, or ignore `ev` entirely
reset_fee=Decimal("99"),
),
)
The Monte-Carlo horizon and the subscription share one unit — a billing month — so
"P(pass)" and "what it costs to find out" are about the same attempt. The honest headline
is breakeven_pass_value: how much a pass must be worth for the attempt to break even.
It requires no assumption about funded-phase performance, which ev does.
Gate: the breakeven pass value is comfortable against your honest estimate of what a funded account is worth to you. If it is uncomfortable, no EV arithmetic rescues it.
The loop, and the caveat that never detaches¶
Failures at stage 5 loop back to stage 2 with a specific fix; a tune at stage 6 loops back to stage 4. The workflow converges when a configuration passes every gate without further edits — at which point you hold: a mechanically-verified strategy, a pass probability with a defensible confidence band, agreement between two opposite-biased estimators, a search that survived its own accounting, and a price.
What you still do not hold is calibration. Every figure rests on rule and fee constants that are researched, cited config — not verified against a live account. The provenance line on every report says so, and it stays true through every stage above. Treat the final answer as a strong diagnostic, size your commitment to that, and let a real blotter close the gap before you treat it as ground truth.