Skip to content

topstep-backtest — agent guide

How to WRITE A CORRECT STRATEGY here, in one file sized to be read in full. Modifying the framework is docs/DESIGN.md; rule sources docs/topstep-rules.md; is-X-built-yet docs/ROADMAP.md; indicator reference docs/INDICATORS.md.

1. What this is

An event-driven backtester answering ONE question: can a strategy profitably pass the Topstep Trading Combine? The rule engine is real-time, not a post-hoc scorecard — the two-state trailing MLL, optional DLL, consistency target, position cap and 16:10 ET flatten all mutate the trade sequence (forced liquidation at adverse slippage plus a $10/contract auto-liquidation fee), so a strategy that would have recovered by Friday can be dead on Tuesday. A strategy reaches the venue only through structural protocols (protocols.py) that BOTH the deterministic SimBroker and the live topstep-sdk AsyncTopstepClient satisfy, so the same class runs in both worlds unchanged; SDK models are imported verbatim.

Scope: Combine only (Funded/XFA parked, docs/topstep-rules.md §6) and Tier-0 fills only (OHLCV bars). One contract → subclass SymbolStrategy; many → subclass Strategy.

2. Quickstart — one complete strategy

Runs as written; the shape of nearly every strategy you will write here.

from __future__ import annotations

from datetime import date
from decimal import Decimal

from topstep_backtest import AccountSize, Backtest, SymbolStrategy
from topstep_backtest.core.instruments import spec_for_symbol
from topstep_backtest.data.synthetic import synthetic_bars
from topstep_backtest.indicators import Cross, Ema
from topstep_backtest.protocols import Bar


class EmaCross(SymbolStrategy):
    """Long on a fast/slow EMA cross-up; exit on the bracket or the cross-down."""

    def __init__(self, contract_id: str, *, fast: int = 12, slow: int = 26) -> None:
        super().__init__(contract_id)
        # Every named indicator IS a TA-Lib function: Ema(12) == TalibIndicator("EMA",
        # timeperiod=12). Nothing here reimplements a formula. Registration order IS
        # update order: use() a Cross's inputs BEFORE the Cross.
        self.fast = self.use(Ema(fast))
        self.slow = self.use(Ema(slow))
        # Cross is the one exception — TA-Lib has no crossover function, so this
        # compares two TA-Lib outputs rather than computing an indicator.
        self.cross = self.use(Cross(self.fast, self.slow))

    async def on_bar(self, bar: Bar) -> None:
        # Fires only for self.contract_id, only once every use()d indicator is ready, each
        # already updated for THIS bar. An order placed here is eligible from the NEXT bar;
        # a market order fills at that next bar's OPEN.
        if self.cross.up and self.position.flat and not self.working_orders:
            await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80)
        elif self.cross.down and self.position.is_long:
            await self.close()


contract = "CON.F.US.MNQ.U26"
bars = synthetic_bars(
    contract_id=contract,
    spec=spec_for_symbol("MNQ"),
    start_day=date(2026, 5, 4),
    days=6,
    seed=7,
    start_price=Decimal("18000.00"),
    bars_per_day=120,
    vol_ticks=12,
)
report = Backtest(bars, EmaCross(contract), account=AccountSize.S50K).run()
print(report)  # verdict, balance path, day trail, summary, provenance
print(report.result.verdict.name)  # Verdict is an IntEnum — print .name

Backtest takes an INSTANCE, never the class; run() is sync, arun() the coroutine. Engine, broker, kernel and strategy state are single-use — fresh Backtest AND fresh strategy per run.

Managing a trade after entry: the bracket is two REAL reduce-only orders (a stop and a limit, OCO-paired) created when the entry fills, so move_stop(ticks=0) / move_target(price=...) amend them in flight — ticks is signed in the POSITION'S favour (ticks=0 is breakeven either way), measured off position.avg_price (the venue's own average, None when flat), and both return how many orders moved with rejections routed to on_reject. They act on bracket children only (parent_order_id set), so a stop-ENTRY of your own is never mistaken for protection. An amended level is live from the NEXT bar.

Real data: Backtest.from_dataframe(df, strategy, *, contract_id, stamp, unit, unit_number) — the spec is derived from contract_id — or build bars yourself with bars_from_dataframe / bars_from_records, which do take an explicit spec. Read §5.3 on stamp first.

To watch a run bar by bar, add record=True: the report is byte-identical (recording is observation only, pinned by a golden), Report.replay carries every frame, and report.to_html(path) grows a second tab holding the replay cockpit — decisions with the broker's answers, rejections included, fills, indicator values named by their attributes, running stats on the tape and beside it, and the session enforcement between bars. self.note("why") inside a hook is the narrative channel: the recorder captures WHAT on its own; only the strategy can say WHY. It is a pure sink — a no-op unrecorded, and never able to influence the run. report.replay_json(path) dumps the raw recording.

3. Public API surface

Generated from the live objects — every signature below is real.

Generated by scripts/gen_api_surface.py — do not hand-edit. Regenerate after changing any public signature; CI fails if it is stale.

Entry point

Name Signature What it does
Backtest (data: 'Sequence[Bar]', strategy: 'Strategy', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, validate: 'bool' = True, record: 'bool' = False, account_id: 'int' = 1, fill_config: 'BarFillConfig \| None' = None, broker_config: 'SimBrokerConfig \| None' = None, fee_model: 'TopstepFees \| None' = None, **rejected: 'object') The two-line runner: assemble the sim stack correctly and run it once.
Report (result: 'BacktestResult', stats: 'SummaryStats', trades: 'tuple[HalfTradeModel, ...]', params: 'CombineParams', bars_gated: 'int \| None' = None, bars: 'tuple[Bar, ...]' = (), instruments: 'dict[str, InstrumentSpec] \| None' = None, replay: 'Replay \| None' = None) One run's full report: the untouched frozen BacktestResult, derived
AccountSize AccountSize.S50K \| AccountSize.S100K \| AccountSize.S150K The three Trading Combine account sizes Topstep offers.
DataValidationError (issues: 'tuple[ValidationIssue, ...]') Bar data failed validation; issues carries every ERROR finding.

Strategy base classes

Name Signature What it does
Strategy (*args, **kwargs) Base strategy. Override the hooks you need; all are optional.
StrategyContext (orders: 'OrderApi', positions: 'PositionApi', history: 'HistoryApi', clock: 'Clock', account_id: 'int', instruments: 'Mapping[str, InstrumentSpec]' = <factory>) Everything a strategy may touch. Protocol-typed: sim and live inject

Strategy (recommended)

Name Signature What it does
SymbolStrategy (contract_id: 'str', require_ready: 'bool' = True, warmup: 'int \| None' = None, trade_sessions: 'Sequence[Session] \| None' = None) Base for strategies trading exactly one contract.

Position / order views

Name Signature What it does
NetPosition (net: 'int' = 0, avg_price: 'Decimal \| None' = None) Signed net contracts for one instrument (positive = long), and the average price they were entered at.
PositionTracker () Per-contract NetPositions, routed by contract_id.
OrderTracker () Latest OrderModel per order id, folded from on_order events.

Data in

Name Signature What it does
bars_from_dataframe (df: 'Any', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]' Build Bar objects from a pandas DataFrame of OHLCV candles.
bars_from_records (rows: 'Iterable[tuple[object, ...]]', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]' Build tick-grid-validated Bar objects from (ts, o, h, l, c, v) rows.

Data in (Parquet export)

Name Signature What it does
load_bars (path: 'str \| Path') -> 'tuple[tuple[Bar, ...], InstrumentSpec, dict[str, Any]]' Read the export -> (bars, spec, metadata), ready for Backtest(...).

Data in (synthetic)

Name Signature What it does
synthetic_bars (*, contract_id: 'str', spec: 'InstrumentSpec', start_day: 'date', days: 'int', seed: 'int', start_price: 'Decimal', bars_per_day: 'int \| None' = None, unit: 'AggregateBarUnit' = <AggregateBarUnit.MINUTE: 2>, unit_number: 'int' = 1, drift_ticks_per_day: 'int' = 0, vol_ticks: 'int' = 8, hours: 'Hours' = 'rth') -> 'tuple[Bar, ...]' Generate days trading sessions of consistent, on-grid OHLCV bars.

Instruments

Name Signature What it does
spec_for_symbol (symbol: 'str') -> 'InstrumentSpec' Look up the built-in spec for a product symbol (raises KeyError if unknown).
symbol_of_contract_id (contract_id: 'str') -> 'str' Extract the product symbol from a gateway contract id.
InstrumentSpec (*args, **kwargs) Frozen per-product economics and session metadata.

Sessions

Name Signature What it does
Session (*args, **kwargs) A named intraday window, defined in its own local timezone.
ASIA Session(name='ASIA', tz='Asia/Tokyo', start=09:00:00, end=15:00:00) A named intraday window, defined in its own local timezone.
LONDON Session(name='LONDON', tz='Europe/London', start=08:00:00, end=16:30:00) A named intraday window, defined in its own local timezone.
NEW_YORK Session(name='NEW_YORK', tz='America/New_York', start=09:30:00, end=16:00:00) A named intraday window, defined in its own local timezone.

Results

Name Signature What it does
BacktestResult (*args, **kwargs) End-of-run outcome (msgspec-serializable -> golden-master friendly).

| SummaryStats | (*args, **kwargs) | Combine-centric summary metrics for one finished backtest run. |

Results (HTML tearsheet)

Name Signature What it does
render_html (report: 'Report', *, replay: 'ReplaySpec' = 'auto', confidence: 'MonteCarloConfidence \| None' = None, crosscheck: 'CrossCheck \| None' = None) -> 'str' Render report as one self-contained interactive HTML document.
render_sweep_html (sweep: 'WindowSweep \| SpacedSweep') -> 'str' Render a sweep as one self-contained HTML document.

Replay recording

Name Signature What it does
Replay (*args, **kwargs) One recorded run: everything the tearsheet's replay scrubber shows.
StatsSnapshot (*args, **kwargs) Running statistics as of a frame's settle — a full SummaryStats over the run's prefix, computed by the sa…
OrderIntent (*args, **kwargs) One strategy decision at the ctx seam, with its outcome.
Recorder () Engine-side run recorder. Observes; never influences.

Tuning

Name Signature What it does
BarFillConfig (*args, **kwargs) Knobs of the Tier-0 pessimism model (all default to the conservative side).

| SimBrokerConfig | (forced_liq_slippage_ticks: 'int' = 2, liquidation_fee_per_contract: 'Decimal' = Decimal('10'), max_trail_ticks: 'int' = 1000, history_depth: 'int' = 20000) | Tunables that are broker-level (not fill-model-level). |

| TopstepFees | (overrides: 'dict[str, FeeSchedule] \| None' = None) | Per-side Topstep fee model (protocols.FeeModel). |

Monte-Carlo

Name Signature What it does
monte_carlo (result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0) -> 'MonteCarloResult' Block-bootstrap paths synthetic Combines from a finished backtest.
MonteCarloResult (*args, **kwargs) Outcome distribution over paths synthetic Combine attempts.
FailureMode FailureMode.MLL_BREACH \| FailureMode.CONSISTENCY_BLOCKED \| FailureMode.TARGET_NOT_REACHED Why a simulated path did not pass. Ordered by when it is decided.

Monte-Carlo confidence

Name Signature What it does
mc_confidence (result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200, lengths: 'Sequence[int]' = (1, 5, 10, 20)) -> 'MonteCarloConfidence' One call: the estimate plus its CI, sensitivity row, and year strata.
MonteCarloConfidence (*args, **kwargs) The point estimate and every qualifier this module can attach to it.
pass_probability_ci (result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200) -> 'PassProbabilityCI' Double bootstrap: a confidence band for the pass probability.
PassProbabilityCI (*args, **kwargs) A pass probability with the error bar its sample size actually earns.
block_length_sensitivity (result: 'BacktestResult', *, params: 'CombineParams', lengths: 'Sequence[int]' = (1, 5, 10, 20), paths: 'int' = 1000, horizon_days: 'int \| None' = None, seed: 'int' = 0) -> 'BlockLengthSensitivity' Re-run the Monte Carlo across block lengths and report the swing.
BlockLengthSensitivity (*args, **kwargs) The same estimate at several block lengths, plus how far it moved.
monte_carlo_by_year (result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 1000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0) -> 'YearStratification' One Monte Carlo per calendar year of the source run.
YearStratification (*args, **kwargs) Per-year estimates, ascending by year.
crosscheck (mc: 'MonteCarloResult', sweep: 'WindowSweep') -> 'CrossCheck' Compare a Monte Carlo against a window sweep, outcome by outcome.
CrossCheck (*args, **kwargs) The bootstrap and the window sweep, forced to answer side by side.

Sequential Combines

Name Signature What it does
sequential_combines (bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'WindowSweep' Run one fresh Combine per non-overlapping window_days-day window.
spaced_combines (bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', periods: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'SpacedSweep' Run periods fresh Combines with start days spread evenly over the tape.
WindowSweep (*args, **kwargs) Every window's attempt, plus the rates over them.
SpacedSweep (*args, **kwargs) A requested number of periods, start days spread evenly, overlap allowed.
WindowResult (*args, **kwargs) One window's Combine attempt, resolved on its own merits.

Overfitting guards

Name Signature What it does
deflated_sharpe (daily_pnl: 'Sequence[Decimal]', *, trials: 'int', trial_sharpe_variance: 'float \| None' = None) -> 'DeflatedSharpe \| None' Deflate a strategy's Sharpe by the number of trials it was selected from.
DeflatedSharpe (*args, **kwargs) A Sharpe ratio and what it is worth once the search is accounted for.
TrialLedger (path: 'str \| Path') A file-backed count of every strategy configuration you have tried.
probability_of_backtest_overfitting (matrix: 'Sequence[Sequence[float]]', *, splits: 'int' = 16) -> 'PBOResult' Run CSCV over a trials x observations performance matrix.
PBOResult (*args, **kwargs) Probability of Backtest Overfitting, by combinatorially symmetric cross-validation (Bailey, Borwein, Lopez de…
sharpe_ratio (daily_pnl: 'Sequence[Decimal]') -> 'float \| None' Per-day Sharpe of a daily P&L series. NOT annualized.

Search & walk-forward

Name Signature What it does
optimize (bars: 'Sequence[Bar]', factory: 'StrategyFactory', grid: 'Sequence[Mapping[str, object]]', *, account: 'AccountSize' = <AccountSize.S50K: '50K'>, objective: 'Callable[[Report], Decimal]' = net_pnl_objective) -> 'OptimizationResult' Run every configuration in grid over bars and keep them all.
OptimizationResult (*args, **kwargs) Every configuration tried, plus the winner — in that order of emphasis.
walk_forward (bars: 'Sequence[Bar]', factory: 'StrategyFactory', grid: 'Sequence[Mapping[str, object]]', *, folds: 'int' = 4, account: 'AccountSize' = <AccountSize.S50K: '50K'>, objective: 'Callable[[Report], Decimal]' = net_pnl_objective) -> 'WalkForwardResult' Anchored walk-forward: select on everything before a fold, score on it.
WalkForwardResult (*args, **kwargs) Anchored walk-forward: each fold re-selects on all data before it.
FoldResult (*args, **kwargs) One walk-forward step: chosen in-sample, scored out-of-sample.
net_pnl_objective (report: 'Report') -> 'Decimal' Default objective: net P&L after all fees.

Attempt economics

Name Signature What it does
evaluate_ev (mc: 'MonteCarloResult', *, economics: 'EvalEconomics') -> 'EvalEV' Combine a bootstrapped outcome distribution with your own prices.
EvalEV (*args, **kwargs) Expected value of one attempt, and the number that needs no assumption.
EvalEconomics (*args, **kwargs) What one evaluation attempt costs YOU. No defaults on the prices.

Continuous contracts

Name Signature What it does
stitch_continuous (series: 'Mapping[str, Sequence[Bar]]', *, spec: 'InstrumentSpec', symbol: 'str \| None' = None, roll_days: 'Mapping[str, date] \| None' = None) -> 'ContinuousSeries' Back-adjust several expiries into one continuous series.
ContinuousSeries (*args, **kwargs) A stitched, back-adjusted series ready to hand to Backtest.
RollEvent (*args, **kwargs) One seam: the day the front month changed, and what it cost to align.

SymbolStrategy methods (what you call inside on_bar)

Name Signature What it does
SymbolStrategy.use (self, indicator: 'T', *, session: 'Session \| None' = None) -> 'T' Register an indicator: auto-updated on every matching bar and
SymbolStrategy.buy (self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None' Buy this contract (market unless a price kwarg implies otherwise);
SymbolStrategy.sell (self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None' Sell this contract (market unless a price kwarg implies otherwise);
SymbolStrategy.close (self) -> 'None' Flatten this contract's position; a rejection goes to on_reject.
SymbolStrategy.cancel_working (self) -> 'None' Cancel every working order on this contract, one cancel per order;
SymbolStrategy.move_stop (self, *, price: 'Decimal \| None' = None, ticks: 'int \| None' = None) -> 'int' Move every bracket stop on this contract; returns how many moved.
SymbolStrategy.move_target (self, *, price: 'Decimal \| None' = None, ticks: 'int \| None' = None) -> 'int' Move every bracket take-profit on this contract; returns how many moved.
SymbolStrategy.note (self, text: 'str') -> 'None' Attach a free-text breadcrumb to the current bar's replay frame.
SymbolStrategy.on_bar (self, bar: 'Bar') -> 'None' —
SymbolStrategy.on_fill (self, trade: 'HalfTradeModel') -> 'None' —
SymbolStrategy.on_reject (self, error: 'APIError') -> 'None' Called with the APIError when a sugar order call is rejected.

Indicators — every one wraps TA-Lib; none is reimplemented here.

Name Signature
Adx (period: 'int' = 14, history: 'int \| None' = None)
Atr (period: 'int', history: 'int \| None' = None)
BBands (period: 'int' = 20, deviations: 'float' = 2.0, history: 'int \| None' = None)
Cross (a: 'ValueSource', b: 'ValueSource')
Ema (period: 'int', history: 'int \| None' = None)
Highest (period: 'int', history: 'int \| None' = None)
Indicator (*args, **kwargs)
Lowest (period: 'int', history: 'int \| None' = None)
Macd (fast: 'int' = 12, slow: 'int' = 26, signal: 'int' = 9, history: 'int \| None' = None)
NotReadyError (*args, **kwargs)
Obv (history: 'int \| None' = None)
Rsi (period: 'int', history: 'int \| None' = None)
Sma (period: 'int', history: 'int \| None' = None)
StdDev (period: 'int', history: 'int \| None' = None)
Stoch (fastk: 'int' = 5, slowk: 'int' = 3, slowd: 'int' = 3, history: 'int \| None' = None)
TalibIndicator (name: 'str', price: 'str \| None' = None, history: 'int \| None' = None, **params: 'int \| float')
TalibLine (owner: 'TalibIndicator', name: 'str')
ValueSource (*args, **kwargs)
WarmIndicator (*args, **kwargs)

BacktestResult fields — the frozen run outcome.

Field Type
verdict Verdict
reason str
ending_balance Decimal
starting_balance Decimal
profit_target Decimal
floor Decimal
best_day Decimal
total_profit Decimal
days_traded int
day_records tuple[DayRecord, ...]
breach Breach | None
trade_count int
equity_curve tuple[tuple[int, Decimal], ...]
rejections tuple[tuple[int, int], ...]
bar_equity tuple[BarEquity, ...]
round_trips tuple[RoundTrip, ...]

SummaryStats fields — derived metrics. See the basis notes below.

Field Type
verdict Verdict
closed_trades int
win_rate Decimal | None
expectancy Decimal | None
profit_factor Decimal | None
max_drawdown Decimal
equity_peak Decimal
final_balance Decimal
net_pnl Decimal
distance_to_floor Decimal
consistency_headroom Decimal
days_traded int
exposure Decimal | None
start_ts_ns int | None
end_ts_ns int | None
provisional bool
avg_win Decimal | None
avg_loss Decimal | None
payoff_ratio Decimal | None
expectancy_r Decimal | None
longest_losing_streak int
breakeven_cost_per_half_turn Decimal | None
sortino Decimal | None
calmar Decimal | None
daily DailyStats
drawdown DrawdownStats
round_trips RoundTripStats

4. The ctx seam

A strategy touches the venue through self.ctx and nothing else — never import a concrete broker, clock or fill model. That is what makes the class run live unchanged.

ctx.orders place / buy / sell / modify / cancel / cancel_all / search_open / get / wait_for_fill — the exact SDK OrderResource signatures, incl. stop_loss_ticks/take_profit_ticks magnitudes and PlaceOrderBracket. Async; raise APIError on refusal.
ctx.positions search_open / close / partial_close / close_all
ctx.history retrieve_bars(cid, *, unit, unit_number, start_time, end_time, limit=1000, live=False, include_partial_bar=False) — serves only already-seen bars; non-native bar specs refused.
ctx.clock now_ns() / now(). Never datetime.now().
ctx.account_id first positional argument to every orders/positions call
ctx.instrument(cid) InstrumentSpec — tick size, tick value, session metadata

SymbolStrategy is sugar over exactly that: buy/sell/close/cancel_working, plus move_stop/move_target over the bracket children (stop_orders/target_orders are the filtered views), self.position (flat/is_long/is_short/avg_price), self.working_orders, self.spec, self.bars_gated. The sugar does not raise — it catches APIError, routes it to on_reject and returns None; self.ctx.orders is the raising path. Both are counted in the rejection tally (§5.5).

Hooks, all optional: on_start (sync), on_bar, on_order, on_fill, on_position, on_stop (sync). Override on_*; drivers call handle_*, so framework bookkeeping cannot be severed. The sim emits no PositionModel event on a FULL close (live sends a size-0 snapshot) — detect flatness from self.position.flat in on_fill, not by overriding on_position.

5. Things that silently produce WRONG answers

Every item here is a bug that runs clean and returns a number.

5.1 Execution

  • Market orders fill at the NEXT bar's open, positions.close/partial_close included. An order participates in a bar only if accepted_ts <= bar.ts_event (its OPEN). That is the no-look-ahead firewall, not latency modelling, and it is not configurable.
  • wait_for_fill raises in sim. React in on_order/on_fill — parity-safe in both worlds.
  • Limit orders need the bar to trade THROUGH the level, ≥1 tick past. An exact touch, including at the open, does not fill unless BarFillConfig(fill_limit_on_touch=True) — a sensitivity probe, not a fix for "it should have filled". Stops gap-fill at the open when the bar opens beyond them, else trigger ± stop_slippage_ticks (default 1) adverse.
  • A bar holding both your stop and your target resolves adverse-extreme-first: the stop wins. One pessimistic O→H→L→C path orders everything by TRIGGER level, never by slippage-adjusted price; breach ties beat fills; equity is re-checked at each fill as it applies.
  • Bracket children are created at the entry FILL (offsets from the ACTUAL fill price) and go active only the bar AFTER it — one bar where the entry is on and its protection is not. Reduce-only orders clamp to the live position and can never flip exposure.
  • OrderModel.trail_price is the trail DISTANCE as a price offset, not the stop level. STOP_LIMIT and JOIN_BID/JOIN_ASK are rejected at Tier-0.
  • APIError.error_code: 4 position cap / account dead / day-locked, 5 no-trade window 16:10–18:00 ET or weekend, 8 unknown contract, 2 validation. They happen live too.

5.2 Indicators

Rules here; tables, per-function warmups and the refused-function list in docs/INDICATORS.md.

  • Never hand-write an indicator. Every value comes from TA-Lib, so no formula can drift. 152 of its 161 functions are reachable via TalibIndicator("NAME", …); the typed wrappers (Sma, Ema, Rsi, Atr, Macd, BBands, Stoch, Adx, Obv, StdDev, Highest, Lowest) are the same machinery with names. Nine functions are REFUSED at construction rather than failing silently mid-run (EXP/COSH/SINH/ACOS/ASIN, MAVP, MAXINDEX/MININDEX/MINMAXINDEX); so are lossy parameters (Sma(14.7), matype=2.9, bools) and per-function-invalid periods (Rsi(1) raises at construction, not 500 bars in).
  • use() it or it never updates, and registration order IS update order — inputs before the Cross reading them. A Cross that was never use()d RAISES from up/down once both inputs are ready; it used to read False forever and take zero trades in silence.
  • Register the OWNER, not a TalibLine. macd.line("macdsignal") has no update on purpose and use() on one raises; use() the Macd, then use(Cross(macd.line("macd"), macd.line("macdsignal"))). A Cross over a Cross is refused (ValueError), a non-Indicator at registration (TypeError) rather than hours in.
  • ready is NOT warm. ready = a value exists (at lookback) and is the use() gate. warm = the bounded buffer is FULL (at history_bars), so the value no longer depends on where this run started. Parity begins at warm: preload history_bars bars live, not lookback. Never write "parity holds once ready". Read history_bars off the instance — usually max(512, 64 × lookback), but SAR() derives it from acceleration (→ 2000 on a lookback of 2), MAMA from slowlimit, KAMA has a 9000-bar floor.
  • Values are Decimal but deliberately NOT tick-snapped — an indicator level is not a tradeable price. Never route one into grid math without an explicit round_to_tick. They are float64-precise, not Decimal-exact: deterministic across reruns, but a Cross on a Bollinger edge can flip on STDDEV's cancellation noise. A band touch is not exact.
  • use(ind, session=…) scopes the DATA; trade_sessions= scopes the DECISION. Two independent switches — see §5.11. Scoping an indicator does not restrict trading, and restricting trading does not starve an indicator.
  • One indicator instance per thread; a parameter sweep gets one per worker.

5.3 Data

  • stamp has no default and never will. It declares what your source timestamp MEANS: stamp="open" → ts_event = ts, ts_init = ts + step; stamp="close" → ts_init = ts, ts_event = ts - step. Getting it wrong is a silent one-bar look-ahead that inflates every result and raises nothing. Confirm your vendor's convention before you pass it.
  • Naive timestamps, off-grid prices and DAY-unit wrangling (the 23h Globex day needs a session-aware resampler that does not exist) are rejected with the fix named. validate=True is the default: any ERROR finding raises DataValidationError carrying issues, and turning it off does not disable the feed's time-order assert.
  • There is NO exchange holiday calendar, by decision — a hand-maintained one was wrong on ~five dates a year and its cleaner ran by DEFAULT, silently deleting tradable sessions. So a holiday bar is indistinguishable from any weekday bar, the broker rejects weekends only, and synthetic_bars(days=N) counts WEEKDAYS, so a tape spanning Thanksgiving or Good Friday emits a session real data would not contain. Filter exchange holidays upstream, in the data you feed in.

5.4 Rules — and the caveat that outranks the rest

  • The MLL breach number is FRAMEWORK-COMPUTED and cannot be reconciled against the gateway. The check runs continuously on realized plus unrealized equity, marked at the bar's adverse extreme (open long at the low, open short at the high) so an intrabar breach is never missed. The real SDK PositionModel has no unrealized_pnl and TradingAccountModel has no equity or open-P&L field — there is no live number to diff it against. The value deciding pass/fail is computed here and nowhere else. Report every verdict as a diagnostic, never as an authoritative pass/fail; the rule and fee constants are cited config, uncalibrated against a real account (docs/topstep-rules.md §9). This is the largest parity hazard in the product.
  • MLL is two-state: the floor ratchets ONLY on end-of-day closed balance and locks permanently at the starting balance once EOD ≥ start + buffer, while breach checks run continuously. Intraday-trailing is Apex, not Topstep — do not port that intuition. A session close at or below the floor is itself a breach; flatten fees alone can do it.
  • The DLL (dll_enabled=True, off by default) is NOT a fail: flatten plus a day-lock cleared at 18:00 ET, no re-emission while locked, and it does not preclude passing.
  • Consistency: best_day <= 0.5 * total_profit. Formula and per-size dollar table are both canonical; no cross-size ratio holds. Negative report.stats.consistency_headroom means a passing balance still fails the Combine.
  • The 16:10 ET flatten is forced liquidation at forced_liq_slippage_ticks adverse plus $10 per contract. Close your own positions first if you care about the price.

5.5 A zero-trade result with a REJECTED line is a wiring bug

Refused placements are tallied, never silent: every order path funnels through one choke point counting rejections by error_code into SimBroker.rejections and BacktestResult.rejections (sorted (error_code, count) pairs), and print(report) emits a REJECTED line whenever it is non-empty. A strategy whose every order was refused — cap 4, outside-hours 5 — used to render as a clean zero-trade report. Always read that line: zero trades WITH rejections is broken wiring; zero trades without is a quiet strategy. Same for report.bars_gated — if it equals your bar count, your warmup is longer than your data.

5.6 Metrics carry a BASIS, and mixing bases is how a report lies

Every figure in SummaryStats states its basis in its docstring because most of them admit two honest answers. The ones that bite:

  • Gross vs net. win_rate, expectancy, expectancy_r, payoff_ratio, avg_win, avg_loss and profit_factor are GROSS; net_pnl is net of every fee. Fees are charged on every half-turn, so a strategy can print profit_factor 1.14 and still lose money — the bundled sma_cross example does exactly that. Never quote a gross ratio beside net_pnl as though they share a basis.
  • Half-turns, not round trips. closed_trades counts CLOSING HALF-TURNS. A flip is one half-turn that closes and reopens. breakeven_cost_per_half_turn is per half-turn for the same reason — that is how fees actually land. There is no round-trip pairing: it needs a FIFO convention whose live-gateway equivalence is unverified (docs/topstep-rules.md §9 box 5).
  • Two R-multiples, and they are not the same number. stats.round_trips.expectancy_r is the TRUE one: net P&L over the dollars actually risked at entry (bracket stop_loss_ticks x tick value x size), per flat-to-flat excursion. stats.expectancy_r is a gross-basis approximation using R = average loss, kept for strategies that set no stops. Prefer the round-trip one when round_trips.with_known_risk > 0; the gap between that and round_trips.count is how much of the R statistics to distrust. Risk is captured AT ENTRY — trailing the stop later does not change it.
  • round_trips are NET of fees; the half-turn figures are GROSS. The basis flips between the two blocks deliberately, because a round trip is a complete decision and the honest question is what it earned after costs. The bundled sma_cross example prints gross profit_factor 1.14 and net round-trip expectancy -3.81 on the same run. The net one is the one that pays you. Counts differ too: a three-clip scale-out is three closed_trades but one round trip.
  • round_trips.worst_r below -1 means a stop was jumped — a gap, or the Tier-0 model's stop slippage. Read it before trusting any risk-based sizing built on "I only ever lose 1R".
  • round_trips.worst_trade is the survival question; worst_r is the strategy question. R says how an excursion went against its own plan, dollars say whether the ACCOUNT could absorb it. A single trip that loses more than the Daily Loss Limit ends the day whatever its R looked like. Read worst_trade against the DLL and against distance_to_floor before either ratio.
  • round_trips.avg_duration_ns / max_duration_ns are nanoseconds, and both endpoints are fill stamps — Tier-0 stamps a fill at its bar's OPEN or CLOSE and never between, so an excursion opened and closed inside one bar reports 0, not its true sub-bar life. Holding time is a rule question here: the engine flattens at 16:10 ET, so a typical duration approaching the session's remainder means the flatten, not the strategy, is closing the trades.
  • Three drawdowns, deliberately. drawdown.static (fixed initial balance), drawdown.eod_trailing (Topstep's actual MLL mechanic) and drawdown.intraday_trailing (Apex-style, ratchets on unrealized highs) are different numbers on the same path and a strategy can survive one while violating another. The legacy max_drawdown is a fourth, close-sampled measure. Say which one you mean.
  • drawdown.avg_eod_trailing averages EPISODES, not days — each excursion below the high-water mark measured at its own trough, over eod_episodes of them, with eod_trailing as the max of the same set. The gap between mean and max is the shape of the risk: a max far above the mean is one bad week, a max close to it means the curve lives at that depth and the worst case is the normal case.
  • equity_peak is close-basis and seeded at the starting balance, so it never reports below it and equity_peak - max_drawdown is the trough that produced it. It matters because a trailing MLL floor is ANCHORED to a peak — but the real one ratchets on end-of-day closed balances, so the floor actually in force followed the EOD peak, which is at or below this. Where the distinction bites, quote drawdown.eod_trailing.
  • exposure is a FRACTION of bars, not a percentage, and it is what tells you how to read every other figure: identical drawdowns at 0.05 and 0.95 exposure are not the same risk. A bar counts when a trip was open strictly inside it, boundary resolved forward (a position opened at a bar's close belongs to the next bar). Multi-instrument runs count a timestamp ONCE — time with risk on, not a sum of per-symbol exposures. None without bars, which is not "never exposed". The denominator includes bars outside tradable hours.
  • sortino and calmar are NOT annualized — daily-dollar and combine-horizon bases respectively. Annualizing a twenty-day sample produces a number with no defensible meaning.
  • daily.stdev is dollars and population-basis (divides by closed_days), matching sortino's downside deviation and likewise not annualized. Dollars because the DLL is a fixed dollar threshold, so this is the dispersion figure that makes "could I breach it" answerable. It is exactly 0 for a one-day run — an artifact of n=1, not a finding.
  • start_ts_ns / end_ts_ns / duration_ns are CALENDAR time, spanning every night and weekend the market was shut. Judge a run's length by days_traded; the two diverge sharply on any sparse or gapped feed.
  • provisional is True below PROVISIONAL_TRADE_FLOOR (200) closes and every rate, ratio and percentile is then an estimate swamped by sampling error. Nothing is suppressed — it is labelled, and print(report) says so loudly. Do not quote a provisional metric as a finding.
  • drawdown.intraday_trailing and min_floor_headroom need BacktestResult.bar_equity, and are None without it (hand-built results, older runs). They are sampled over the modelled Tier-0 intrabar path, not real ticks — conservative estimates of a tick-based figure, not reproductions of one. min_floor_headroom is the closest the account ever came to termination; distance_to_floor is only where it ENDED.

5.7 The Monte-Carlo answers a different question, and has its own traps

metrics.monte_carlo(result, params=..., paths=..., seed=...) block-bootstraps the run's observed trading days and replays each synthetic sequence through a fresh CombineKernel.

  • The kernel decides every verdict. Nothing in that module re-implements the MLL ratchet, the consistency test or the pass condition — it drives the same three calls the live engine uses. Never add a second rulebook there.
  • Read the autopsy, not the pass probability. mll_breach / consistency_blocked / target_not_reached partition the failures and imply different fixes: resize, throttle the outsized day, or accept the edge is too slow. The headline number alone tells you nothing actionable.
  • block_length=1 is a footgun. It degenerates to an i.i.d. resample, destroys the losing streaks that actually blow accounts, and will report a pass probability that is far too kind. Default is 5.
  • horizon_days defaults to BILLING_MONTH_DAYS (21), never the observed day count. A Combine has no time limit, only a monthly fee, so "one attempt" defaults to one fee cycle — the same unit EvalEconomics bills in. The fixed default is also the safety property: a blown run stops recording days at the breach, so its count is a survival time, not a Combine length, and is never used as a horizon. source_truncated still flags that such a sample is survivorship-biased by construction — the days after the blow-up do not exist, every figure is conditioned on having survived, and nothing can repair that; do not quote such a result without the caveat.
  • provisional below 30 source days. Resampling cannot create information.
  • The point estimate ships with its own cross-examination (metrics/confidence.py). pass_probability_ci double-bootstraps the source days themselves — the error bar the day count earns, which the path count never was. block_length_sensitivity shows whether the streak assumption is load-bearing. monte_carlo_by_year refuses to average a hostile year against a kind one. crosscheck compares against sequential_combines under a binomial 2×SE null — when they disagree, the disagreement is the finding. Quote a pass probability with its CI, not alone; mc_confidence bundles the lot and report.to_html(path, confidence=..., crosscheck=...) renders the cards.
  • It cannot invent a regime the tape never contained, and it inherits every uncalibrated constant (§5.4). A probability to three decimals from unverified inputs is precise, not accurate.

5.8 One long backtest answers a question nobody is asked

Nobody runs a single continuous evaluation for three years. They attempt a Combine, and if it resolves, they attempt another. metrics/windows.py::sequential_combines models that: the tape is cut into consecutive non-overlapping windows of window_days trading days, each run as a completely fresh evaluation (new balance, new floor, new broker, new strategy instance).

  • It is the empirical twin of the Monte-Carlo, biased the opposite way. monte_carlo resamples observed days — many paths, real sequencing destroyed beyond the block length. This preserves sequencing and regime exactly and pays in sample size: three years is ~35 independent 21-day windows, not 35,000. They share classify_failure so the autopsies are directly comparable. Run both; disagreement is the finding.
  • Non-overlapping is not a limitation, it is the point. Stepping one day at a time would give ~700 windows from three years, each sharing 20 of 21 days with its neighbour — the precision of 700 observations carrying the information of 35. The step is fixed at the window length rather than shipped as a footgun with a warning.
  • pass_rate is an estimate off a handful of samples. At ~35 attempts its standard error is near 8 percentage points before anything else is considered. Always read it beside attempts.
  • Warm starts happen OUTSIDE the engine, and must. Preload bars fed through a Backtest would land in day_records as flat days and move closed_days, every daily percentile, stdev, sortino and the drawdown durations while leaving P&L untouched. So SymbolStrategy.prewarm drives the preceding history through the indicators only — no on_bar, no orders — and the run then sees only its own window. prewarm raises if the strategy is already bound.
  • history_bars, not warmup, is the preload size (§5.2's ready vs warm). That is real money: Sma(30) is warm at max(512, 64 x 30) = 1920 bars, which on RTH-only 5-minute candles is ~25 trading days of history per window. Expect the first windows of any tape to be fully_warm=False — cold windows UNDER-trade, so leaving them in biases pass_rate DOWN.
  • trailing_days_dropped is reported, never absorbed. A partial window is not an attempt, and scoring one would count a short evaluation as a failure to reach the target.

5.9 Optimizing is a search, and a search lies by default

optimize() was absent from this project on purpose — maximizing over a Combine metric is an overfitting machine. It exists now because the guards do. Use it accordingly:

  • OptimizationResult.best is not evidence. Selecting a maximum from a search guarantees a flattering number. The API keeps every trial specifically so you can feed pbo_matrix to probability_of_backtest_overfitting and the winner's daily_pnl to deflated_sharpe. A bare best-params answer is the thing to refuse.
  • PBO and walk-forward are not substitutes. PBO asks whether your SELECTION generalises and is symmetric in time, so it says nothing about decay. Walk-forward asks whether the edge SURVIVES FORWARD and is directional, so it conflates a decaying edge with a bad selection rule. Run both; they disagree in informative ways.
  • Read efficiency, not the OOS number. How much of the in-sample edge survived is the question. Negative efficiency — chosen for making money, then lost money — is the signature of a fit to noise. It is None against a non-positive in-sample objective, because a ratio against a loss inverts the sign of good and bad.
  • consistent_folds guards the aggregate. A good total built from one huge fold and three losers is not a robust strategy.
  • The default objective is net P&L, deliberately not the verdict. Optimizing a binary pass/fail throws away nearly all the information in a run and rewards configurations that scraped over the line once.
  • Cost: a grid of G over F folds runs G x (F + 1) complete backtests. Serial by design — a search that silently sampled would be worse than a slow one.

5.10 Multi-year runs need stitching, and the seam corrupts INDICATORS not P&L

A two-year run spans ~8 quarterly rolls. data/continuous.py::stitch_continuous turns per-expiry bars into one back-adjusted series labelled with the bare ticker ("MNQ", which spec_for_symbol already resolves), so the engine sees one instrument and one unbroken account.

  • A raw splice does NOT produce wrong P&L. The 16:10 flatten plus the session-roll backstop mean no position and no working order survives a day boundary, and a roll seam IS a day boundary — so entry and exit always share one contract. What the seam corrupts is indicator state, which does span it: a spurious Cross, an Atr spike inflated for a whole lookback, a false Highest/Lowest breakout. The resulting trades are priced correctly and should never have been taken.
  • Additive back-adjustment is therefore exactly P&L-neutral, because entry and exit share the offset and it cancels in the difference. Pinned end to end by test_engine_pnl_is_invariant_under_a_constant_price_shift — if that ever fails, additive stitching is unsafe and this reasoning is wrong.
  • Never ratio-adjust. Multiplicative offsets do not cancel (P&L would move) and push prices off the tick grid, which SimBroker asserts on every fill.
  • What adjustment does distort: logic keyed to ABSOLUTE price levels (round numbers, a fixed price target). Tick-relative logic — stop_loss_ticks, take_profit_ticks, all the strategy sugar — is unaffected. Adjusted prices also no longer match live history.retrieve_bars, which returns raw single-expiry quotes.
  • The roll spread is measured at ONE shared instant from both expiries, so the two contracts must overlap; a spread taken across a time gap folds that interval's market move into the adjustment permanently. Rolls are forced onto day boundaries — that is what makes the cancellation argument hold.
  • Still your problem on a multi-year tape: exchange holidays (~20/year, no calendar ships — filter upstream) and the uncalibrated constants (§5.4).

5.11 Sessions scope indicator DATA and trading DECISIONS, separately

ASIA / LONDON / NEW_YORK live in core/sessions.py; the worked example is examples/session_scoped.py.

  • The two switches are independent, and conflating them is the bug. use(Atr(14), session=NEW_YORK) restricts which bars that indicator is computed from. SymbolStrategy(..., trade_sessions=(NEW_YORK,)) restricts when on_bar may fire. Indicators advance regardless of trade_sessions — an indicator fed only the tradable window develops gaps and computes a different value from the same tape. Both default to today's behaviour, and an unscoped strategy is byte-identical to one written before sessions existed.
  • Which indicators to scope is a modelling decision, and it is not uniform. Dispersion measures (Atr, StdDev, Rsi, Stoch, BBands) describe how much price moves PER BAR, and that is session-dependent: on a 24h feed an Atr(14) read at 09:30 ET is computed almost entirely from thin pre-market bars, so it understates NY volatility exactly when stop distance is being sized. Level measures (Sma, Ema) answer where price IS, and the overnight move is real — an NY-only Ema is anchored to yesterday's 16:00 close. Levels are continuous across sessions; dispersion is not.
  • A scoped indicator warms in ITS OWN cadence. It needs history_bars bars of its session, so on 5-minute bars an NY-scoped Sma(30) spans ~25 trading days against ~7 unscoped. Read strategy.warm (per-indicator update counts) or history_bars_by_session — never bars_seen >= history_bars, which cannot express two cadences and OVERSTATES warmth. sequential_combines still SIZES its preload slice from the unscoped history_bars, so a scoped strategy will honestly report fully_warm=False rather than silently lying.
  • A Cross inherits its inputs' scope and refuses a conflicting session=. Mixed-scope inputs are refused outright: they advance on different bars, so comparing them compares values sampled at unrelated instants.
  • Sessions are defined in their own timezone, not as fixed ET offsets. DST comes from the IANA database, so nothing rots — London and New York switch on different dates, and the London window really is 04:00 ET rather than 03:00 for ~3 weeks each spring and ~1 each autumn. A fixed ET block is still one line: Session("LONDON_ET", ET, time(3), time(11)).
  • Session membership is a DIFFERENT axis from trading_day_of(). Asia sits after the 18:00 ET rollover, so its bars belong to the NEXT trading day. Membership is tested on ts_event (the bar's OPEN) over a half-open [start, end) window, so the 09:29→09:30 bar — every print of it pre-market — is not New York.
  • trade_sessions only narrows what THIS strategy does. It never widens what the venue permits: the 16:10 ET flatten and the 16:10–18:00 no-trade window apply either way.
  • You need 24h data. The shipped data/sample_mnq_1m.csv is RTH-only (390 bars/day, 09:30–15:59 ET) and contains no Asia or London bars, so nothing session-scoped is observable on it. For synthetic bars pass synthetic_bars(..., hours="globex"); the default "rth" mode is the New York session and nothing else.
  • Not built: per-session performance attribution. Nothing in SummaryStats splits P&L by session, so scoping is currently a modelling choice you make, not one the report scores.

6. The end-to-end workflow

What using this framework actually looks like, in the order you do it:

from topstep_backtest import AccountSize, Backtest
from topstep_backtest.metrics import crosscheck, mc_confidence, sequential_combines
from topstep_backtest.rules.params import combine_params

# 1. Write the strategy (§2), then run ONE backtest — recorded, so the replay
#    tab can show intent against execution.
report = Backtest(bars, MyStrategy(CONTRACT), account=AccountSize.S50K, record=True).run()
print(report)  # verdict, day trail, four statistics blocks

# 2. Sanity-check the wiring BEFORE reading any number.
assert not report.result.rejections  # zero trades + rejections = broken wiring (§5.5)
assert report.bars_gated != len(bars)  # warmup longer than your data

# 3. Read the run, minding the basis (§5.6) — and step the replay tab to each
#    entry: the bracket where you meant it, the notes matching the fills.
s = report.stats
if s.provisional:
    ...  # under 200 closes: estimates, not findings
s.round_trips.expectancy  # NET per trade — the one that pays you
s.round_trips.expectancy_r  # true R, when with_known_risk > 0
s.drawdown.eod_trailing  # Topstep's actual MLL mechanic
s.drawdown.min_floor_headroom  # closest the account came to death
s.daily.p05  # the bad day to size against

# 4. Stop trusting one sample — real attempts first, resampled ones second.
sweep = sequential_combines(bars, lambda: MyStrategy(CONTRACT), window_days=21)
c = mc_confidence(report.result, params=combine_params(AccountSize.S50K), seed=7)
check = crosscheck(c.mc, sweep)  # same billing-month horizon on both sides (§5.7)

# 5. Act on the AUTOPSY, and quote the CI, never the bare point (§5.7).
#    mll_breach          -> resize
#    consistency_blocked -> throttle the outsized day; the edge is fine
#    target_not_reached  -> the edge is too slow; nothing risk-side helps
c.ci.p05, c.ci.p95  # the error bar the source-day count earns
c.sensitivity.spread  # wide = the streak assumption is doing the work
check.divergent  # bootstrap vs real windows: disagreement is the finding
report.to_html("run.html", confidence=c, crosscheck=check)  # the cards, archived

# 6. Was it an edge, or did you search until something looked good? The trial
#    count must be RECORDED — a remembered one is always too low, because the
#    sweeps you abandoned are exactly the ones that inflated the winner.
ledger = TrialLedger(".trials.json")
ledger.record("sma-10-30", sharpe=sharpe_ratio(daily))
dsr = deflated_sharpe(daily, trials=ledger.count(), trial_sharpe_variance=ledger.sharpe_variance())
dsr.survives  # deflated >= 0.95, the conventional bar
dsr.expected_max_sharpe > dsr.sharpe  # the search alone explains the result

# 7. Is the attempt worth its price? Every figure is yours; none is baked in.
ev = evaluate_ev(
    c.mc,
    economics=EvalEconomics(
        monthly_fee=Decimal("149"), pass_value=Decimal("2000"), reset_fee=Decimal("99")
    ),
)
ev.breakeven_pass_value  # lead with THIS — it assumes nothing

Steps 2 and 3 are not optional politeness. A zero-trade run with rejections and a gross-basis ratio quoted as if it were net are the two most common ways to read a confidently wrong result out of this framework.

Steps 6 and 7 have one trap each. DSR with trials=1 deflates nothing — reporting it is how a user reveals they never counted the search, and the statistic exists precisely to stop that. EV's pass_value is an assumption you supply and this package cannot check; it dominates the answer, so quote breakeven_pass_value ("a pass must be worth at least $X") rather than ev unless you can defend the input.

Runnable end to end: examples/run_combine.py (steps 1–3), examples/run_windows.py and examples/run_montecarlo.py (steps 4–5, the latter at two position sizes so the autopsy visibly discriminates). docs/TUTORIAL_EMA_CROSSOVER.md walks all of it line by line, and website/workflow.md is the narrative version with the gate each stage must pass.

7. What Backtest() refuses, and what to use instead

Coming from backtesting.py, the dialect is deliberately close but causal. Backtest(bars, s, cash=50_000, commission=0.002) raises ValueError: Backtest() does not support '<knob>': … naming the alternative. Verbatim, from harness.py::_REJECTED_KNOBS:

Knob The error's explanation
cash= account economics are fixed by the Combine — pick account=AccountSize.S50K/S100K/S150K instead
commission= fees are the real Topstep schedule (fills.fees.TopstepFees), applied always; they are not a per-trade knob. To correct a rate against your own blotter, pass the whole schedule: fee_model=TopstepFees(overrides={...})
margin= there is no margin model — the Combine's MLL floor and position cap (combine_params) are the risk limits
spread= execution costs come from the Tier-0 bar fill model (fills.bar_fill.BarFillConfig slippage ticks), not a synthetic spread
trade_on_close= filling at the decided bar's close is the exact look-ahead the accepted_ts firewall forbids; orders fill from the NEXT bar
hedging= the venue nets positions per contract — hedged positions cannot exist
exclusive_orders= hidden auto-close orders would mutate the strategy's intent sequence (the parity gate); make reversals explicit awaits
finalize_trades= end-of-session positions are governed by the 16:10 ET flatten rule, never a stats toggle

Also refused: the strategy class instead of an instance (TypeError naming the fix); fractional or equity-fraction sizing (size is int contracts under a hard cap); init()-time full-array precompute (self.I), since there is no full array at the live edge; and synchronous order placement — the SDK is async and the awaits ARE the live contract. Translations:

backtesting.py here
init() / next() __init__ / async def on_bar(self, bar)
self.I(SMA, self.data.Close, n) self.fast = self.use(Sma(n))
arrays lose their NaNs the ready-gate; count in report.bars_gated, pin with warmup=
self.data.Close[-1] the bar argument; ctx.history.retrieve_bars(…) for seen history
crossover(fast, slow) self.use(Cross(fast, slow)) → .up / .down
self.buy(size=2, sl=…, tp=…) await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80)
trade.sl = price / trade.tp = price await self.move_stop(price=…) / move_target(price=…) — or ticks= from position.avg_price, signed in the position's favour. These amend REAL working orders, so a rejection is real too
stats = bt.run() → 30-key Series report = bt.run() → Report + .stats/.result/.trades
optimize() metrics.optimize(...) — but read §5.9 first; it keeps every trial on purpose
plot() report.show() — interactive HTML tearsheet (report.to_html(path) to name the file; bt.run_with_tearsheet(dir) to get one from every run; Backtest(..., record=True) adds the bar-by-bar replay tab)

8. Knobs — the calibration seam

Four keyword arguments, and only these, move economics or pessimism.

Backtest(
    bars,
    strategy,
    account=AccountSize.S100K,  # 50K/100K/150K — balance, floor, target, position cap
    dll_enabled=True,  # optional Daily Loss Limit (off by default)
    fill_config=BarFillConfig(  # topstep_backtest.fills.bar_fill
        stop_slippage_ticks=1,  # adverse ticks on EVERY triggered stop (gap or intrabar)
        market_slippage_ticks=0,  # adverse ticks on market fills
        fill_limit_on_touch=False,  # True = optimistic touch fills (a sensitivity probe)
    ),
    broker_config=SimBrokerConfig(  # topstep_backtest.execution.sim_broker
        forced_liq_slippage_ticks=2,  # MLL/DLL/16:10 market-dump slippage
        liquidation_fee_per_contract=Decimal("10"),  # Topstep's auto-liquidation fee
        max_trail_ticks=1000,
        history_depth=20000,
    ),
    fee_model=TopstepFees(
        overrides={
            "MNQ": FeeSchedule(  # topstep_backtest.fills.fees
                commission_per_side=Decimal("0.25"),
                exchange_per_side=Decimal("0.37"),
                nfa_per_side=Decimal("0.02"),
            )
        }
    ),
)

fee_model= is a calibration seam, not a P&L dial: the shipped rates are researched but uncalibrated, and this is how you substitute what your own blotter charges. It is the only fee input; per-trade commission= is refused (§7). Every default sits on the conservative side — moving one toward optimism is a probe to report next to the baseline, not a new baseline.

9. Checklist before returning a strategy

  1. Subclasses SymbolStrategy/Strategy; reaches the venue only via self.ctx or the sugar; no concrete broker, no datetime.now().
  2. Every indicator is self.use()d — inputs before their Cross, owners not lines.
  3. on_bar/on_order/on_fill are async def; every order call is awaited.
  4. Entry guarded on self.position.flat and not self.working_orders.
  5. Size is an int number of contracts, within the account cap.
  6. Ran end to end once and you read the report — verdict, REJECTED line, bars_gated.
  7. report.result.rejections is empty, or every code in it is one you intended.
  8. On real data: stamp= set from the vendor's documented convention, holidays filtered upstream.
  9. The verdict is reported as a diagnostic, not a pass/fail (§5.4).

10. Map: module → what it owns

Module Owns
protocols.py Clock / Bar / DataFeed / Broker / OrderApi / PositionApi / HistoryApi / FillModel / FeeModel / Fill / WorkingOrder
core/ money (tick grid, P&L), instruments (15-product spec table), time (ET sessions), ids
rules/ CombineKernel — THE rulebook state machine, golden-fixtured; params per account size
clock/ deterministic TestClock, wall-clock LiveClock
fills/ Tier-0 BarFillModel/BarFillConfig, shared build_path, TopstepFees/FeeSchedule
data/ wrangler (bars_from_dataframe/bars_from_records, explicit stamp), validator, synthetic, feed (ListBarFeed), clean (infer_symbol_from_filename/AmbiguousSymbolError — filename→symbol inference when loading CSVs)
execution/ SimBroker/SimBrokerConfig (lifecycle, OCO brackets, netting, breach walk, forced liq), rejections
engine/ BacktestEngine loop, BacktestResult (msgspec-serializable, carries rejections)
strategy/ Strategy + StrategyContext (the write-once seam), SymbolStrategy (use(), ready-gate, sugar), tracker (NetPosition/PositionTracker/OrderTracker, pure folds over SDK models)
metrics/stats.py SummaryStats + compute_summary — every derived figure off one run; DrawdownStats, DailyStats, RoundTripStats. Bases in §5.6
metrics/montecarlo.py monte_carlo — pass probability + violation autopsy, replayed through the REAL CombineKernel, never a second rulebook (§5.7)
metrics/overfitting.py deflated_sharpe, TrialLedger, probability_of_backtest_overfitting (§5.9). Float, not Decimal — normal-theory statistics
metrics/windows.py sequential_combines — one long tape replayed as consecutive independent Combine attempts, warm-started off the preceding history (§5.8)
metrics/walkforward.py optimize (keeps EVERY trial) + anchored walk_forward + split_by_day (§5.9)
metrics/economics.py evaluate_ev + EvalEconomics — EV per attempt from YOUR prices; lead with breakeven_pass_value
data/continuous.py stitch_continuous — many expiries → one back-adjusted series on the bare ticker (§5.10)
indicators/ TA-Lib adapter (152 of 161 functions), typed wrappers, Cross, Indicator/ValueSource
replay.py Recorder + Replay — bar-by-bar run recording behind Backtest(record=True): decisions at the ctx seam (via recording proxies), events, indicator values, state tracks, sparse SummaryStats snapshots built by the SAME metrics/stats helpers (terminal snapshot pinned equal to compute_summary). Observation only: results are byte-identical either way
harness.py Backtest facade (one shared clock, instruments from the feed, strict validation, refuses backtesting.py knobs) and Report

11. Not built — do not write code against it

docs/ROADMAP.md is the authority. These names appear in older notes and other frameworks and do not exist here: LiveBroker, RecordingLiveBroker, DataEngine, TrailingMaxLossLimit, prob_fill_on_limit (the real field is BarFillConfig.fill_limit_on_touch). Also absent: a Parquet/Arrow data catalog, multi-timeframe resampling, a MessageBus, per-year / per-regime / per-session breakdowns (§5.11), higher fill tiers, an XFA rule set, and any holiday calendar.

Do write code against these — they used to be on the list above and now ship: Monte-Carlo (§5.7), optimize() and the overfitting guards (§5.9), the sequential-Combine sweep (§5.8), the continuous-contract stitcher (§5.10), and the HTML tearsheet (report.show() / report.to_html(path) / Backtest.run_with_tearsheet(dir), which writes a timestamped file per run; print(report) remains the text render). Their signatures are in the generated table in §3.

Parity status, precisely: structural conformance is proven — a pyright-strict protocol conformance test plus a signature diff of SimOrderApi.place against the SDK's OrderResource.place (tests/parity/test_broker_conformance.py), both run in CI. The behavioural gate — an intent-sequence test asserting an identical Submit/Modify/Cancel sequence sim vs live — is Phase 4 and is not met. Behavioural parity is argued, not proven.

12. Modifying the framework?

Read docs/DESIGN.md first: it carries the architecture contract this file assumes — Decimal on the tick grid with FIFO position accounting, int-ns UTC time with ET boundaries, the accepted_ts no-look-ahead firewall, intrabar determinism along one pessimistic path, and the rule that protocols.py is THE canonical interface module (change a signature there first or not at all).

The gate is seven commands, and all seven run in CI (.github/workflows/ci.yml, on Python 3.12/3.13/3.14, plus a job that installs the built wheel into a clean venv and smoke-tests it and a job that builds the documentation site):

uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv run python scripts/gen_api_surface.py --check   # §3 is stale
uv run python scripts/check_doc_links.py           # a doc link dangles
uv run mkdocs build --strict                       # needs --extra docs

ruff is pinned >=0.16,<0.17 on purpose — format output changes between minors, so a floating pin turns ruff format --check red on an unrelated upgrade. pyproject.toml declares exactly three extras, [data], [dev] and [docs]. If you change the public surface, regenerate §3 with uv run python scripts/gen_api_surface.py — and add the new names to GROUPS in that script, or they will be documented nowhere and the --check will still pass.

The documentation site

uv run --extra docs mkdocs serve renders it locally at http://127.0.0.1:8000. mkdocs.yml sets docs_dir: website, which holds ONLY the site-original pages (index, quickstart, results, beyond-one-backtest). Everything else is assembled at build time by scripts/gen_docs.py:

  • The repo's own markdown (README.md, this file, CHANGELOG.md, docs/*.md) and the runnable examples are mirrored into the build at their repo-relative paths. That layout is load-bearing — docs/DESIGN.md links to ../AGENTS.md and the tutorial links to ../examples/ema_cross.py, and mirroring keeps ONE link graph that works in a checkout, in the sdist and on the site. Never "fix" a link for the site alone; you would fork the graph and rot whichever copy nobody reads.
  • The API reference is one mkdocstrings stub per module, listed in MODULES in that script. A new public module needs an entry there and a nav: entry in mkdocs.yml; miss the first and it is documented nowhere, miss the second and the strict build fails.

show_if_no_docstring: true is REQUIRED in the mkdocstrings options, not cosmetic: this codebase shares one docstring across groups of related fields (p05..p95, best_r/worst_r), griffe attaches it to the LAST name only, and the default would silently drop every other member.