Skip to content

The workflow

A backtest number is cheap to produce and expensive to trust. This framework's tools are not a menu — they are a sequence, and each stage exists because the stage before it produces a number that looks finished but is not. The order below is the order in which questions become answerable: run mechanics before verdicts, distributions before probabilities, guards before tuning, prices before decisions.

Each stage ends with a gate — the question you must be able to answer before the next stage's output means anything. The gates are where the discipline lives; the code is short.

# Stage Tool Gate
0 Know the envelope operating limits Is this strategy even measurable here?
1 Data in, declared Backtest.from_dataframe Validator clean, stamp known from vendor docs
2 Write the strategy SymbolStrategy Brackets at entry, TA-Lib only, warmup understood
3 One recorded run record=True + tearsheet The strategy did what you meant — verdict ignored
4 Real attempts sequential_combines Enough attempts to say anything; failure mix known
5 The distribution mc_confidence + crosscheck CI narrow enough to act on; no unexplained divergence
6 Tuning (only now) optimize + guards DSR survives the honest trial count; PBO low
7 Price the attempt evaluate_ev Breakeven pass value is comfortable, not heroic

Stage 0 — Know the envelope before writing code

Some strategies cannot be measured by this engine, and the failure is silent: the run completes, the report is plausible, the number is wrong. Check yours against the operating limits first:

  • One contract per run. Feed a single contract's bars (already-continuous data is fine; stitch_continuous exists if yours is not).
  • The edge must clear the fill model's pessimism. Tier-0 fills are deliberately harsh — next-bar-open market fills, trade-through limits, stop slippage. A per-trade edge under ~8–10 ticks is dominated by that pessimism, not measured by it. Tick-scalping and limit-heavy passive entries are out of scope.
  • Holidays and half-sessions are YOUR problem. No calendar ships, on purpose; filter them upstream or they are scored as ordinary days.
  • Every rule and fee constant is cited config, not calibration (§9). Every verdict downstream of here is a diagnostic.

Gate: the strategy's edge, instruments and session behavior fit inside the envelope. If not, no downstream number — however good — is about your strategy.

Stage 1 — Data in, declared and validated

bt = Backtest.from_dataframe(
    df,
    MyStrategy(CONTRACT),
    contract_id=CONTRACT,
    stamp="open",  # from the VENDOR'S docs, never from a trial run
    unit=AggregateBarUnit.MINUTE,
    unit_number=5,
)

stamp, unit and unit_number have no defaults because guessing the stamp is a silent one-bar look-ahead that inflates every result, and a wrong stamp on RTH-only data is byte-identically silent. Validation is strict by default — let it refuse. And check the contract id twice: a wrong --symbol validates cleanly and is ~25× wrong (wrong tick value, no error anywhere).

Gate: the validator passes, and you know the stamp from documentation rather than from which setting produced nicer results.

Stage 2 — Write the strategy against the dialect

Subclass SymbolStrategy (quickstart, tutorial). Three habits pay for themselves downstream:

  • Register every indicator with use() — TA-Lib only, updated once per bar, warmup gated. No hand-written formulas; none can drift from the reference implementation.
  • Attach the bracket at entry (stop_loss_ticks= / take_profit_ticks=). A stop at entry is what defines the trade's initial risk, and initial risk is what makes every R-multiple in the report a true R rather than a retrofitted one.
  • Narrate with self.note(). Notes are pure observation — they cannot change a result — and they are what makes stage 3's replay legible: the tape shows what happened, the notes show what you thought was happening.

Gate: you can state the strategy's intent per bar without reference to hidden state, and you know its warmup cost (history_bars — real money at ~25 trading days of 5-minute RTH history for a 30-bar indicator).

Stage 3 — One recorded run, read for mechanics, not verdict

report = Backtest(bars, MyStrategy(CONTRACT), record=True).run()
print(report)
report.to_html("run.html")  # results tab + the replay · bar by bar tab

The single run's job is not to tell you whether the strategy passes. It is one sample — acting on it is how a strategy that got lucky once becomes a funded account that blows up. Its job is to prove the mechanics:

  • REJECTED count is zero or every rejection is explained. A strategy whose orders were refused reports a clean-looking run that measured nothing.
  • Warmup gating matches expectation — bars_gated should equal your indicator math, not surprise you.
  • The replay tab shows intent matching execution: step to each entry, check the bracket is where you meant it, watch the notes against the fills.
  • The fills are believable against how each metric is based.

Gate: the strategy demonstrably did what you meant. Only now is a verdict worth computing — on more than one sample.

Stage 4 — Stop trusting one sample: real attempts

from topstep_backtest.metrics import sequential_combines

sweep = sequential_combines(bars, lambda: MyStrategy(CONTRACT), window_days=21)

Sequential combines cut the tape into consecutive fresh evaluations — real sequencing, real regimes, prewarmed indicators. Expect the head of the tape fully_warm=False, and expect TARGET_NOT_REACHED to dominate: that is the ordinary outcome of a 21-day window, not a failure signal. Read pass_rate beside attempts — three years is ~35 windows, a standard error near 8 points before anything else is considered.

Gate: you have enough attempts for the rate to mean anything, and you know which way the failing windows failed.

Stage 5 — The distribution, quoted with its confidence

from topstep_backtest.metrics import crosscheck, mc_confidence

c = mc_confidence(report.result, params=report.params, seed=7)
check = crosscheck(c.mc, sweep)  # horizons must match: both one billing month
report.to_html("run.html", confidence=c, crosscheck=check)

The Monte Carlo turns your observed days into thousands of synthetic attempts through the real rule kernel; the confidence instruments say how much the resulting number is worth. Read them in this order:

  1. The autopsy before the probability — mll_breach / consistency_blocked / target_not_reached imply different fixes (resize / throttle the outsized day / better edge), and the fix loops you back to stage 2, not to a bigger position.
  2. The CI before the point. A 65% from 500 source days and a 65% from 40 are different findings. Quote the band.
  3. The block-length spread — wide means the streak assumption is doing the work, and the honest report is the range.
  4. The year strata — the pooled number averages a hostile year against a kind one; the spread across years is the error bar non-stationarity imposes.
  5. The crosscheck — the bootstrap and the real windows estimate the same quantity with opposite biases. When they disagree, the disagreement is the finding.

Gate: the CI is narrow enough to act on, nothing is flagged provisional, and the crosscheck shows no unexplained divergence. If the CI spans "great" to "hopeless", the answer is a longer tape, not a decision.

Stage 6 — Tuning, only now, and only with the guards on

If — and only if — the mechanics are proven and the distribution says the edge is real but mis-sized or mis-tuned:

from topstep_backtest.metrics import optimize, walk_forward
from topstep_backtest.metrics import (
    TrialLedger,
    deflated_sharpe,
    probability_of_backtest_overfitting,
)

optimize keeps every trial precisely so the guards can price the search: feed the trial matrix to PBO, the winner's daily P&L and the honest trial count (the TrialLedger, which remembers the sweeps you abandoned) to deflated Sharpe, and confirm with anchored walk_forward that the edge survives forward. Then re-run stages 4–5 on the tuned configuration — a tuned strategy is a new strategy.

Gate: DSR survives at the recorded trial count and PBO is low. A best-params answer without those two numbers is the thing to refuse — including from yourself.

Stage 7 — Price the attempt

from topstep_backtest.metrics import EvalEconomics, evaluate_ev

ev = evaluate_ev(
    c.mc,
    economics=EvalEconomics(
        monthly_fee=Decimal("165"),  # YOUR prices — nothing is baked in
        pass_value=Decimal("2500"),  # your honest estimate, or ignore `ev` entirely
        reset_fee=Decimal("99"),
    ),
)

The Monte-Carlo horizon and the subscription share one unit — a billing month — so "P(pass)" and "what it costs to find out" are about the same attempt. The honest headline is breakeven_pass_value: how much a pass must be worth for the attempt to break even. It requires no assumption about funded-phase performance, which ev does.

Gate: the breakeven pass value is comfortable against your honest estimate of what a funded account is worth to you. If it is uncomfortable, no EV arithmetic rescues it.


The loop, and the caveat that never detaches

Failures at stage 5 loop back to stage 2 with a specific fix; a tune at stage 6 loops back to stage 4. The workflow converges when a configuration passes every gate without further edits — at which point you hold: a mechanically-verified strategy, a pass probability with a defensible confidence band, agreement between two opposite-biased estimators, a search that survived its own accounting, and a price.

What you still do not hold is calibration. Every figure rests on rule and fee constants that are researched, cited config — not verified against a live account. The provenance line on every report says so, and it stays true through every stage above. Treat the final answer as a strong diagnostic, size your commitment to that, and let a real blotter close the gap before you treat it as ground truth.