Skip to content

topstep-backtest — design decision record

An event-driven backtesting framework for developing futures strategies that can profitably pass the Topstep Trading Combine (and survive the Express Funded phase). Takes standardized OHLCV candlestick data first; extends to L1 quotes, L2/MBP-10 depth, and full L3/MBO order-flow without changing strategy code.

Companion docs: topstep-rules.md (the researched rule reference, with sources + confidence) and ROADMAP.md.

What this file is. The decisions behind the framework, the rationale that makes each one load-bearing, the invariants they bind, and the risks they trade against — what someone modifying the framework needs. Not a tutorial and not a file map: AGENTS.md owns using the framework, ROADMAP.md owns "is X built yet". Where this record and the source disagree, the source wins and this file is the bug.


1. Scope

A standalone package whose single most important property is backtest/live parity with topstep-sdk: a strategy is written once against structural protocols and runs unchanged in (a) an in-process deterministic simulator, to prove it can pass a Combine, and (b) live via AsyncTopstepClient, to actually trade it. The differentiator is a first-class prop-firm rule engine — trailing max-loss, daily-loss, consistency, position caps, session flatten — modeled as a real-time risk engine that changes the trade sequence itself, not a post-hoc report on an equity curve.

Current scope: the Trading Combine ONLY. The Express Funded / funded-account (XFA) phase — starting-$0 balance, scaling plan, payout paths, the post-first-payout MLL→0 trap — is explicitly deferred. The design keeps a rule-parameter seam so funded is a later drop-in, but nothing in the roadmap builds it and no Combine work may depend on it; funded rules stay parked in topstep-rules.md §6 as reference. Other non-goals for now: options, cross-venue portfolio margining, sub-microsecond HFT, smart order routing.

2. Standing principles

  1. Parity via swapped edges (the NautilusTrader/LEAN pattern). Keep the strategy fixed; swap only three edges — the Clock, the feed, the Broker. Backtest = TestClock + ListBarFeed + SimBroker; live = LiveClock + a market-hub feed + AsyncTopstepClient. The strategy touches only ctx.orders / ctx.positions / ctx.clock — never a concrete venue or fill model.

  2. Reuse the SDK's domain vocabulary verbatim. OrderModel, PositionModel, HalfTradeModel, PlaceOrderBracket, every enum and APIError are imported from topstep_sdk, never redefined, and SimBroker constructs real SDK model instances for callbacks — so on_order/on_fill/on_position receive the exact classes they get live. That is the strongest available type-parity guarantee, and it is why strategy/tracker.py folds only SDK models.

  3. Fill realism is a hierarchy, not a switch. One FillModel protocol; implementations of increasing fidelity (OHLCV → L1 → L2 → MBO). The strategy imports none of them; the engine owns order lifecycle, OCO/brackets and trailing stops, so every FillModel only ever sees a static resting order. Tier-0 (fills/bar_fill.py) is the only one built.

  4. Determinism is a testable property. Backtests drive the identical async callbacks synchronously from the data iterator — no real concurrency — so msgspec.json.encode(result) is byte-equal across reruns. Any change making a run depend on wall time, dict iteration order, or unseeded randomness is a bug regardless of what else it improves.


3. Decisions and their rationale

3.1 Package boundary — no topstep-core extraction yet

Ship topstep-backtest as a separate package that depends on topstep-sdk and reuses its models verbatim. Do not redefine domain types and do not extract a shared topstep-core yet — defer that until a third consumer (e.g. a live dashboard) appears, then extract with re-exports. Define the parity Protocols in the backtester (the consumer that needs write-once), so the SDK stays a pure API client.

3.2 Three-phase settle per data point

Cascading orders must settle within one instant, exactly as live sequencing does: (1) the broker ingests the point and matches resting orders; (2) the engine dispatches strategy callbacks; (3) the command queue drains and re-matches until empty. A bracket child or hedge submitted from on_fill therefore settles at the same timestamp rather than a bar later. TUTORIAL_EMA_CROSSOVER.md holds the authoritative walkthrough of the current loop.

3.3 Continuous contracts — the adjustment-vs-raw parity hazard

A causal roll stitcher (front month by volume/OI) is the intended shape, and it carries a hazard to design around rather than discover: sim indicators would run on a back-adjusted continuous series while live history.retrieve_bars returns raw single-expiry prices. Reconcile it explicitly — reconstruct the same adjusted series live, or trade raw front-month in sim — and add a roll-boundary test asserting sim and live indicator inputs match. Not built.

3.4 The gateway's P&L lot method must be replicated, not assumed

HalfTradeModel.profit_and_loss is set on the closing half-turn by some lot method (FIFO vs weighted-average) over a netted position. Daily NET is unaffected either way, but per-trade R-multiples and round-turn pairing diverge if the method differs — and those are what a strategy is judged on. This repo uses FIFO lots with an exact cost basis; the obligation is to capture real fills and confirm the gateway agrees, not to trust the choice.

3.5 The fill-tier swap is a falsification test

Running one strategy across tiers on overlapping dates is a diagnostic, not a fidelity upgrade: a large gap between an L2 estimate and an MBO result for a passive/limit-heavy strategy means its edge was a queue-position artifact. Build the tiers so this comparison stays cheap.

3.6 One rule kernel, many callers

The trailing-drawdown / daily-loss / consistency math is implemented once, as rules/kernel.py::CombineKernel — a pure state machine over (day, intraday-equity-path) encoding the two-state trailing MLL, optional DLL, consistency and position cap. Every consumer (sim risk path, any future live shadow-monitor, single-path analytics, a vectorized Monte-Carlo twin) delegates to it or is diff-tested against it. Re-implementation is banned: four divergent copies producing "sim says PASS, analytics says FAIL" is the single most credibility-destroying bug this product can have (§8.1).

3.7 Framework-computed equity is the parity hazard

The breach-triggering number does not come from Topstep. The SDK's PositionModel has no unrealized_pnl and TradingAccountModel has no equity/open-P&L field, so live equity is computed by this framework — which makes it a first-class parity hazard rather than an implementation detail. Compute unrealized in ONE shared place (core/money.py::unrealized_pnl / position_unrealized, division-free, FIFO cost basis) used by the sim broker and by anything that later shadows a live account, pin the mark convention, and calibrate against real liquidations. Two functions computing "equity" is the same failure as two rule kernels.

3.8 Anti-overfitting is domain-mandatory

Optimizing directly on in-sample pass-probability is the dominant failure mode here ("pass 95%, blow up 5%, negative EV") — the Combine's pass/fail framing invites exactly the search that destroys the result. Any sweep must estimate pass-probability on OOS days only, report trial count, PBO (probability of backtest overfitting) and DSR (deflated Sharpe) as primary numbers, gate on a minimum session count, and feed days-to-pass into the EV — the Combine has no time limit but does have a monthly fee, so a strategy that "passes" after eight months of fees is not a win. Report the full outcome distribution and the tail (P(MLL breach), p05 terminal profit, bootstrap CIs), never a single pass rate.

This is now enforced in code, not aspiration. metrics/walkforward.py::optimize retains every trial rather than the winner alone, and TrialLedger carries that count into metrics/overfitting.py::deflated_sharpe; probability_of_backtest_overfitting implements CSCV; anchored walk_forward reports efficiency; metrics/montecarlo.py::monte_carlo block-bootstraps observed days through the real CombineKernel for the distribution and the autopsy; metrics/confidence.py puts the mandated bootstrap CI on that probability (plus block-length sensitivity, per-year strata, and the cross-check against the window sweep); metrics/economics.py::evaluate_ev folds days-to-pass and YOUR prices into an EV and a breakeven. optimize() was withheld until those guards existed — the ordering was the point, and reversing it (shipping a sweep whose output nothing deflates) would re-open this failure mode.

3.9 The replay recorder observes at the engine's seams — never inside them (2026-08-17)

Backtest(record=True) answers "what did the strategy see and decide on THIS bar", and the design constraint that shaped it is that the answer must be worthless-proof: recording can never change the run. Three placements follow:

  • Capture at existing seams, add none. The engine already has the phase boundaries (commit_bar after a bar fully settles, on_session around enforcement, on_user_event in the dispatch loop); decisions are captured by wrapping the strategy's ctx.orders/ ctx.positions in structurally-conformant recording proxies, so raw and sugar order paths are seen identically and the broker is untouched. State tracks (position, working orders) fold the SAME SDK events the strategy's own views fold (strategy/tracker.py) — never broker internals, so the recording shows what a live session would have shown.
  • Running statistics may not become a second implementation. Snapshots call the same metrics/stats.py helpers compute_summary calls (extracted for exactly this), over prefix data; the only hand-rolled parts are the O(1) equity folds (peak/max-DD/static/ intraday trio), and the terminal snapshot is pinned byte-equal to compute_summary by tests/golden/test_replay_goldens.py and swept across tapes by tests/property/test_replay_props.py. Trade figures inside snapshots re-read broker.trades (the ledger), not the recorder's event fold — events lag the ledger inside a session roll, and a snapshot mixing the two would be internally inconsistent.
  • The narrative channel is a pure sink. Strategy.note() writes into the recording and nowhere else; unrecorded it is a no-op. Byte-identity of results with recording on/off and with/without notes is pinned by goldens — if either pin ever breaks, every recording made since is suspect, which is why they are goldens and not unit tests.

The tearsheet embeds the recording behind the same determinism contract as the rest of the render (Python-preformatted strings; JS computes nothing beyond differences of two displayed numbers), windowed with a loud label past REPLAY_AUTO_FRAME_LIMIT frames because a scrubber silently missing five sixths of a run reads as the whole run.

The replay lives in its own tab with its own charts, not inline under the results sheet. Two reasons, and the second is the load-bearing one: a cockpit worth stepping through wants the chart, the state and the event log in one viewport rather than stacked down a scrolling column; and a veil belongs only on a run mid-flight. Sharing one set of charts meant the finished sheet was dimmed by wherever the cursor happened to sit — the results view is the run that finished, and it should say so. The cost is a second chart set built from the same payload (lazily, on first open: a chart created in a display:none container measures zero and opens on an empty time scale) and the drawing helpers shared, so two views of the same bars cannot drift apart.

3.10 Sessions scope DATA and DECISIONS on separate switches, in local time

Two decisions, and each had a tempting wrong answer.

Why two switches rather than one. The obvious design is a single "trade the NY session" flag that filters the tape. It is wrong: an indicator fed only the tradable window develops gaps and computes a different value from the same bars, so the flag would silently change every indicator in the strategy as a side effect of a decision about trading. So use(indicator, session=…) scopes an indicator's INPUT DATA and trade_sessions= scopes when on_bar fires, and neither implies the other. Indicators advance regardless of trade_sessions; a scoped indicator does not restrict trading. That the two are independent is the feature — a continuous trend filter beside a session-scoped volatility measure is the configuration the domain actually wants, and one switch cannot express it.

The domain reason it is worth expressing: dispersion measures (Atr, StdDev, Rsi) describe how much price moves per bar, which is session-dependent — a 24h Atr(14) read at 09:30 ET is computed almost entirely from thin pre-market bars and understates NY volatility exactly when stop distance is being sized. Level measures (Sma, Ema) answer where price is, and the overnight move is real. Levels are continuous across sessions; dispersion is not. The framework stays neutral and makes the choice expressible per indicator.

Why local timezones rather than fixed ET windows. A table of ET windows ("London = 03:00–11:00 ET") is the same mistake as the exchange calendar this project deliberately dropped (§ROADMAP): a hand-maintained mapping that is silently wrong on a predictable schedule. London and New York observe daylight saving but switch on different dates, so such a table is an hour off for roughly three weeks each spring and one each autumn. Defining each session in its own zone (Europe/London, Asia/Tokyo, America/New_York) delegates that to the IANA database, so there is nothing to maintain and nothing to rot. It also removes a branch: in its own timezone no real session wraps midnight. Session remains a plain value type, so a caller who genuinely wants fixed ET blocks writes one — wrapping windows are supported for exactly that case.

Consequence accepted: a scoped indicator warms in its own cadence, needing history_bars bars of its session — roughly 3.5× more calendar time on NY-only 5-minute bars. Warmth is therefore counted per indicator (SymbolStrategy.warm) rather than from a bar total, which cannot express two cadences; sequential_combines asks the strategy instead of comparing counts, because the old preload >= needed test overstates warmth for a scoped indicator.


4. Why the strategy dialect is shaped this way (2026-07-22)

The authoring surface deliberately resembles backtesting.py — one class, a per-bar decision method, declared indicators with automatic warmup gating, position/working-order views, bracketed buy()/sell(), a two-line Backtest(...).run(), a printable report — and stops resembling it at exactly four idioms, refused on constitutional grounds:

  • init() with full-length data + self.I full-array precompute. There is no full array at the live edge, and progressive truncation masks length only: non-causal functions still leak the future. This is the exact leak class the project structurally bans.
  • Sync order placement / sync next(). The SDK is async; orders must be awaited. A layer hiding the awaits inserts a new intent-scheduling machine directly onto the parity seam.
  • trade_on_close. Look-ahead by construction.
  • cash=/commission=/spread=/margin=, fractional-of-equity sizing, hedging, exclusive_orders, finalize_trades. Each is either an equities-world economic knob that would fork sim economics from Topstep's, or a hidden intent mutation that breaks the intent-sequence identity a strategy is supposed to have. harness.py refuses them by name rather than ignoring them, so the mistake surfaces at call time.

Rejected grafts: sync next() with a deferred intent flush (a new parity surface for zero semantic gain); string-keyed Stats (object-typed under pyright strict — SummaryStats is a frozen msgspec struct with typed attributes instead); class-attribute parameter injection (typed constructor params instead); full progressive self.data arrays.

The float-containment amendment (2026-07-24). Adopting TA-Lib meant accepting float64 buffers, whose named cost was price leakage. It is contained, not avoided: floats live ONLY inside the indicator layer; values come back as Decimal and are deliberately never tick-snapped, so an indicator level can never be mistaken for a tradeable price; and every money path reads bar/ctx, never an indicator. That containment rule is binding.


5. Binding invariants (the maintainer contract)

Violations are bugs, not style. Most are pinned by a named test.

  • Money. Decimal on the tick grid, always; all grid math via core/money.py. FIFO lots with an exact cost basis; unrealized through position_unrealized (division-free) — never average-then-divide. Indicator values are the one exception (§4): not tick-snapped, never routed into grid math without an explicit round_to_tick.
  • Time. int-ns UTC in the hot path; canonical ET at boundaries (flatten 16:10, session close 17:00, day reset 18:00). All time flows through the Clock protocol, never datetime.now(). dt_to_ns/ns_to_dt are exact integer math, no float round-trips.
  • No look-ahead, structurally. Bars stamp ts_init at CLOSE; an order participates in a bar only if accepted_ts <= bar.ts_event (its OPEN); trailing stops ratchet only from bars the order lived through; SimHistoryApi serves only already-seen bars; the engine asserts feed time-order. Structural, so they cannot be forgotten — keep them that way.
  • Intrabar determinism. Fills and rule breaches are ordered along ONE pessimistic price path (fills/path.py::build_path, adverse extreme first) by TRIGGER level — never by slippage-adjusted prices; breach ties beat fills; equity is re-checked AT each fill after it applies. Limit fills require trade-THROUGH unless fill_limit_on_touch=True.
  • Rejections. Always the SDK's APIError carrying the gateway's numeric error_code (execution/rejections.py), and always tallied — every order path funnels through one choke point into SimBroker.rejections / BacktestResult.rejections. A strategy whose every order was refused used to render as a clean zero-trade report; it must never do so again.
  • protocols.py is THE canonical interface module and is frozen. Change a signature there first or not at all. tests/parity/test_broker_conformance.py holds a pyright-strict conformance proof plus a keyword-level signature diff of SimOrderApi.place against the SDK's OrderResource.place, so SDK drift fails pytest rather than live trading.
  • _compute clears its dirty flag only AFTER the TA-Lib call succeeds. The order is not stylistic: clearing it first memoises the failure — the exception surfaces once, then every later read that bar reports a benign "not ready", which SymbolStrategy._gated() silently swallows, so a broken indicator becomes a strategy that just never trades. Never reorder those two lines. Pinned by tests/unit/test_talib_adapter_hardening.py::test_a_raising_compute_is_not_cached_as_not_ready.
  • Toolchain gate — all FOUR commands, every time: uv run pytest, uv run ruff check ., uv run ruff format --check ., uv run pyright. .github/workflows/ci.yml runs exactly these on Python 3.12/3.13/3.14, plus a wheel job that builds the distributions, installs the wheel into a clean venv from PyPI only, and runs the documented quickstart outside the source tree — that job is what catches missing package data and imports that work from a checkout but not from an install. release.yml publishes on a v* tag. Three of four is a red build, and ruff format --check is the one people forget — which is also why ruff is pinned >=0.16,<0.17: format output changes between minors, so an unbounded pin would fail CI on correctly formatted code. Runtime deps are topstep-sdk, msgspec, ta-lib (the last is core, not an extra); exactly three extras exist, [data] (pandas + pyarrow, so both DataFrame and Parquet input work), [dev] and [docs].

6. Testing idioms

  • Golden masters for every rulebook worked example (tests/golden/); determinism is msgspec.json.encode(result) equality across reruns.
  • Hypothesis property suites (tests/property/, seeded/derandomized) for money, kernel and indicator invariants: ticks→Decimal→ticks round-trips exactly, the MLL floor is monotone non-decreasing and never exceeds the lock ceiling, every realized trade has an integer tick delta, streaming indicator output equals a batch run bit for bit.
  • Parity conformance in tests/parity/ (§5). Build test bars with tests/conftest.py::minute_bar — ET wall-clock in, exact ns out; a bar "at 9:30" OPENS 9:30 and CLOSES 9:31.

7. Backtest/live parity — how it gets proven, not asserted

  1. Type parity: import + construct real SDK models; on_* callbacks receive identical classes.
  2. Interface parity: Broker/OrderApi/PositionApi copied field-for-field from the SDK resources; a pyright-strict + runtime conformance test asserts AsyncTopstepClient satisfies Broker, plus a signature diff pinned to the SDK version so a mismatch fails pytest, and CI runs that suite on every push and pull request.
  3. Rejection parity: SimBroker raises the SDK's APIError with matching error_code.
  4. Time parity: all clock reads flow through the swappable Clock.
  5. Data parity: the replay feed emits SDK-shaped bars/quotes; a live feed adapts client.market_hub() with the same (contract_id, data) callback shape.
  6. Intent-sequence conformance (the definitive proof): run ONE strategy through SimBroker and through a recording live-broker double, and assert an identical ordered sequence of Submit/Modify/Cancel intents.
  7. Captured-gateway golden fixtures + calibration harness: validate the sim's constructed OrderModel/HalfTradeModel/PositionModel field-by-field against real recorded gateway events, and calibrate fills, fees, latency and liquidations against a real eval account.
  8. Documented divergences you must NOT hide: bar-tier market orders fill next-bar-open in sim but instantly live (a genuine price shift, not merely "conservatism"); stored vendor OHLCV bars are not the tape-built bars live produces. Require Tier-1+ (quote/tick) validation for any strategy whose edge is intrabar-timing-sensitive before trusting it live.

What is actually proven today (0.2.0). Rungs 1–4 hold: the sim constructs real SDK models, and tests/parity/ asserts — pyright-strict and at runtime — that both AsyncTopstepClient and SimBroker satisfy Broker, with a keyword-level signature diff on place. That is structural conformance: same call surface, same types, same APIError codes. Rung 5 is half-built (the replay side ships; there is no live feed). Rungs 6–7 are not built — no recording live-broker double exists anywhere in the repo, and there are no captured-gateway fixtures — so behavioural parity is unproven: nothing yet demonstrates that the same strategy emits the identical Submit/Modify/Cancel sequence against a live broker, or that a constructed HalfTradeModel matches a real one field-by-field. Read the list as the obligation and rungs 6–7 as outstanding work.


8. Ranked risk register

The failures that would silently invalidate a verdict, worst first. Each is designed out, not watched for.

  1. Divergent rule implementations → ONE pure kernel, golden-fixtured, diff-tested against every consumer. (§3.6)
  2. MLL breach on bar data with undefined intrabar ordering → adverse-extreme equity along the single shared price path + deterministic resolution by trigger level. (§5)
  3. Sim equity ≠ Topstep equity (framework-computed unrealized) → one shared unrealized function, pinned mark convention, calibrated against real liquidations. (§3.7)
  4. Forced-liquidation slippage unmodeled → a distinct liquidation path with its own worse slippage (forced_liq_slippage_ticks) plus liquidation_fee_per_contract; realized loss can exceed the floor, decisive at the XFA zero-buffer case.
  5. Market-order fill-timing parity leak (next-bar-open vs instant) → documented, and Tier-1+ validation required for intrabar-sensitive edges. (§7.8)
  6. Overfitting / tail-blindness → full outcome distribution, OOS-only pass-probability, PBO/DSR, stress resamplers. (§3.8)
  7. Cross-subsystem interface conflicts → protocols.py frozen first + conformance test. (§5)
  8. Decimal↔int-tick rounding drift → centralized in core/money.py, property-tested. (§5)
  9. Continuous-contract adjustment vs live raw prices → reconcile + roll-boundary test. (§3.3)
  10. Missing data-quality gate at ingest → data/validator.py::validate_bars (tick grid, duplicates, monotonicity, maintenance-halt and weekend bars, session gaps). Exchange holidays are deliberately not modeled at all — a wrong calendar fabricates a clean-looking result, so filter them upstream, in whatever produces the bar files.