topstep-backtest — agent guide¶
How to WRITE A CORRECT STRATEGY here, in one file sized to be read in full. Modifying the
framework is docs/DESIGN.md; rule sources docs/topstep-rules.md; is-X-built-yet
docs/ROADMAP.md; indicator reference docs/INDICATORS.md.
1. What this is¶
An event-driven backtester answering ONE question: can a strategy profitably pass the Topstep
Trading Combine? The rule engine is real-time, not a post-hoc scorecard — the two-state
trailing MLL, optional DLL, consistency target, position cap and 16:10 ET flatten all mutate
the trade sequence (forced liquidation at adverse slippage plus a $10/contract
auto-liquidation fee), so a strategy that would have recovered by Friday can be dead on Tuesday.
A strategy reaches the venue only through structural protocols (protocols.py) that BOTH the
deterministic SimBroker and the live topstep-sdk AsyncTopstepClient satisfy, so the same
class runs in both worlds unchanged; SDK models are imported verbatim.
Scope: Combine only (Funded/XFA parked, docs/topstep-rules.md §6) and Tier-0 fills only
(OHLCV bars). One contract → subclass SymbolStrategy; many → subclass Strategy.
2. Quickstart — one complete strategy¶
Runs as written; the shape of nearly every strategy you will write here.
from __future__ import annotations
from datetime import date
from decimal import Decimal
from topstep_backtest import AccountSize, Backtest, SymbolStrategy
from topstep_backtest.core.instruments import spec_for_symbol
from topstep_backtest.data.synthetic import synthetic_bars
from topstep_backtest.indicators import Cross, Ema
from topstep_backtest.protocols import Bar
class EmaCross(SymbolStrategy):
"""Long on a fast/slow EMA cross-up; exit on the bracket or the cross-down."""
def __init__(self, contract_id: str, *, fast: int = 12, slow: int = 26) -> None:
super().__init__(contract_id)
# Every named indicator IS a TA-Lib function: Ema(12) == TalibIndicator("EMA",
# timeperiod=12). Nothing here reimplements a formula. Registration order IS
# update order: use() a Cross's inputs BEFORE the Cross.
self.fast = self.use(Ema(fast))
self.slow = self.use(Ema(slow))
# Cross is the one exception — TA-Lib has no crossover function, so this
# compares two TA-Lib outputs rather than computing an indicator.
self.cross = self.use(Cross(self.fast, self.slow))
async def on_bar(self, bar: Bar) -> None:
# Fires only for self.contract_id, only once every use()d indicator is ready, each
# already updated for THIS bar. An order placed here is eligible from the NEXT bar;
# a market order fills at that next bar's OPEN.
if self.cross.up and self.position.flat and not self.working_orders:
await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80)
elif self.cross.down and self.position.is_long:
await self.close()
contract = "CON.F.US.MNQ.U26"
bars = synthetic_bars(
contract_id=contract,
spec=spec_for_symbol("MNQ"),
start_day=date(2026, 5, 4),
days=6,
seed=7,
start_price=Decimal("18000.00"),
bars_per_day=120,
vol_ticks=12,
)
report = Backtest(bars, EmaCross(contract), account=AccountSize.S50K).run()
print(report) # verdict, balance path, day trail, summary, provenance
print(report.result.verdict.name) # Verdict is an IntEnum — print .name
Backtest takes an INSTANCE, never the class; run() is sync, arun() the coroutine. Engine,
broker, kernel and strategy state are single-use — fresh Backtest AND fresh strategy per run.
Managing a trade after entry: the bracket is two REAL reduce-only orders (a stop and a
limit, OCO-paired) created when the entry fills, so move_stop(ticks=0) /
move_target(price=...) amend them in flight — ticks is signed in the POSITION'S favour
(ticks=0 is breakeven either way), measured off position.avg_price (the venue's own
average, None when flat), and both return how many orders moved with rejections routed to
on_reject. They act on bracket children only (parent_order_id set), so a stop-ENTRY of
your own is never mistaken for protection. An amended level is live from the NEXT bar.
Real data: Backtest.from_dataframe(df, strategy, *, contract_id, stamp, unit, unit_number) —
the spec is derived from contract_id — or build bars yourself with bars_from_dataframe /
bars_from_records, which do take an explicit spec. Read §5.3 on stamp first.
To watch a run bar by bar, add record=True: the report is byte-identical (recording is
observation only, pinned by a golden), Report.replay carries every frame, and
report.to_html(path) grows a second tab holding the replay cockpit — decisions with the
broker's answers, rejections included, fills, indicator values named by their attributes,
running stats on the tape and beside it, and the session enforcement between bars. self.note("why") inside a hook is the narrative channel:
the recorder captures WHAT on its own; only the strategy can say WHY. It is a pure sink — a
no-op unrecorded, and never able to influence the run. report.replay_json(path) dumps the
raw recording.
3. Public API surface¶
Generated from the live objects — every signature below is real.
Generated by scripts/gen_api_surface.py — do not hand-edit. Regenerate after changing any public signature; CI fails if it is stale.
Entry point
| Name | Signature | What it does |
|---|---|---|
Backtest |
(data: 'Sequence[Bar]', strategy: 'Strategy', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, validate: 'bool' = True, record: 'bool' = False, account_id: 'int' = 1, fill_config: 'BarFillConfig \| None' = None, broker_config: 'SimBrokerConfig \| None' = None, fee_model: 'TopstepFees \| None' = None, **rejected: 'object') |
The two-line runner: assemble the sim stack correctly and run it once. |
Report |
(result: 'BacktestResult', stats: 'SummaryStats', trades: 'tuple[HalfTradeModel, ...]', params: 'CombineParams', bars_gated: 'int \| None' = None, bars: 'tuple[Bar, ...]' = (), instruments: 'dict[str, InstrumentSpec] \| None' = None, replay: 'Replay \| None' = None) |
One run's full report: the untouched frozen BacktestResult, derived |
AccountSize |
AccountSize.S50K \| AccountSize.S100K \| AccountSize.S150K |
The three Trading Combine account sizes Topstep offers. |
DataValidationError |
(issues: 'tuple[ValidationIssue, ...]') |
Bar data failed validation; issues carries every ERROR finding. |
Strategy base classes
| Name | Signature | What it does |
|---|---|---|
Strategy |
(*args, **kwargs) |
Base strategy. Override the hooks you need; all are optional. |
StrategyContext |
(orders: 'OrderApi', positions: 'PositionApi', history: 'HistoryApi', clock: 'Clock', account_id: 'int', instruments: 'Mapping[str, InstrumentSpec]' = <factory>) |
Everything a strategy may touch. Protocol-typed: sim and live inject |
Strategy (recommended)
| Name | Signature | What it does |
|---|---|---|
SymbolStrategy |
(contract_id: 'str', require_ready: 'bool' = True, warmup: 'int \| None' = None, trade_sessions: 'Sequence[Session] \| None' = None) |
Base for strategies trading exactly one contract. |
Position / order views
| Name | Signature | What it does |
|---|---|---|
NetPosition |
(net: 'int' = 0, avg_price: 'Decimal \| None' = None) |
Signed net contracts for one instrument (positive = long), and the average price they were entered at. |
PositionTracker |
() |
Per-contract NetPositions, routed by contract_id. |
OrderTracker |
() |
Latest OrderModel per order id, folded from on_order events. |
Data in
| Name | Signature | What it does |
|---|---|---|
bars_from_dataframe |
(df: 'Any', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]' |
Build Bar objects from a pandas DataFrame of OHLCV candles. |
bars_from_records |
(rows: 'Iterable[tuple[object, ...]]', *, contract_id: 'str', spec: 'InstrumentSpec', unit: 'AggregateBarUnit', unit_number: 'int', stamp: "Literal['open', 'close']") -> 'tuple[Bar, ...]' |
Build tick-grid-validated Bar objects from (ts, o, h, l, c, v) rows. |
Data in (Parquet export)
| Name | Signature | What it does |
|---|---|---|
load_bars |
(path: 'str \| Path') -> 'tuple[tuple[Bar, ...], InstrumentSpec, dict[str, Any]]' |
Read the export -> (bars, spec, metadata), ready for Backtest(...). |
Data in (synthetic)
| Name | Signature | What it does |
|---|---|---|
synthetic_bars |
(*, contract_id: 'str', spec: 'InstrumentSpec', start_day: 'date', days: 'int', seed: 'int', start_price: 'Decimal', bars_per_day: 'int \| None' = None, unit: 'AggregateBarUnit' = <AggregateBarUnit.MINUTE: 2>, unit_number: 'int' = 1, drift_ticks_per_day: 'int' = 0, vol_ticks: 'int' = 8, hours: 'Hours' = 'rth') -> 'tuple[Bar, ...]' |
Generate days trading sessions of consistent, on-grid OHLCV bars. |
Instruments
| Name | Signature | What it does |
|---|---|---|
spec_for_symbol |
(symbol: 'str') -> 'InstrumentSpec' |
Look up the built-in spec for a product symbol (raises KeyError if unknown). |
symbol_of_contract_id |
(contract_id: 'str') -> 'str' |
Extract the product symbol from a gateway contract id. |
InstrumentSpec |
(*args, **kwargs) |
Frozen per-product economics and session metadata. |
Sessions
| Name | Signature | What it does |
|---|---|---|
Session |
(*args, **kwargs) |
A named intraday window, defined in its own local timezone. |
ASIA |
Session(name='ASIA', tz='Asia/Tokyo', start=09:00:00, end=15:00:00) |
A named intraday window, defined in its own local timezone. |
LONDON |
Session(name='LONDON', tz='Europe/London', start=08:00:00, end=16:30:00) |
A named intraday window, defined in its own local timezone. |
NEW_YORK |
Session(name='NEW_YORK', tz='America/New_York', start=09:30:00, end=16:00:00) |
A named intraday window, defined in its own local timezone. |
Results
| Name | Signature | What it does |
|---|---|---|
BacktestResult |
(*args, **kwargs) |
End-of-run outcome (msgspec-serializable -> golden-master friendly). |
| SummaryStats | (*args, **kwargs) | Combine-centric summary metrics for one finished backtest run. |
Results (HTML tearsheet)
| Name | Signature | What it does |
|---|---|---|
render_html |
(report: 'Report', *, replay: 'ReplaySpec' = 'auto', confidence: 'MonteCarloConfidence \| None' = None, crosscheck: 'CrossCheck \| None' = None) -> 'str' |
Render report as one self-contained interactive HTML document. |
render_sweep_html |
(sweep: 'WindowSweep \| SpacedSweep') -> 'str' |
Render a sweep as one self-contained HTML document. |
Replay recording
| Name | Signature | What it does |
|---|---|---|
Replay |
(*args, **kwargs) |
One recorded run: everything the tearsheet's replay scrubber shows. |
StatsSnapshot |
(*args, **kwargs) |
Running statistics as of a frame's settle — a full SummaryStats over the run's prefix, computed by the sa… |
OrderIntent |
(*args, **kwargs) |
One strategy decision at the ctx seam, with its outcome. |
Recorder |
() |
Engine-side run recorder. Observes; never influences. |
Tuning
| Name | Signature | What it does |
|---|---|---|
BarFillConfig |
(*args, **kwargs) |
Knobs of the Tier-0 pessimism model (all default to the conservative side). |
| SimBrokerConfig | (forced_liq_slippage_ticks: 'int' = 2, liquidation_fee_per_contract: 'Decimal' = Decimal('10'), max_trail_ticks: 'int' = 1000, history_depth: 'int' = 20000) | Tunables that are broker-level (not fill-model-level). |
| TopstepFees | (overrides: 'dict[str, FeeSchedule] \| None' = None) | Per-side Topstep fee model (protocols.FeeModel). |
Monte-Carlo
| Name | Signature | What it does |
|---|---|---|
monte_carlo |
(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0) -> 'MonteCarloResult' |
Block-bootstrap paths synthetic Combines from a finished backtest. |
MonteCarloResult |
(*args, **kwargs) |
Outcome distribution over paths synthetic Combine attempts. |
FailureMode |
FailureMode.MLL_BREACH \| FailureMode.CONSISTENCY_BLOCKED \| FailureMode.TARGET_NOT_REACHED |
Why a simulated path did not pass. Ordered by when it is decided. |
Monte-Carlo confidence
| Name | Signature | What it does |
|---|---|---|
mc_confidence |
(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200, lengths: 'Sequence[int]' = (1, 5, 10, 20)) -> 'MonteCarloConfidence' |
One call: the estimate plus its CI, sensitivity row, and year strata. |
MonteCarloConfidence |
(*args, **kwargs) |
The point estimate and every qualifier this module can attach to it. |
pass_probability_ci |
(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 2000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0, outer: 'int' = 200, inner_paths: 'int' = 200) -> 'PassProbabilityCI' |
Double bootstrap: a confidence band for the pass probability. |
PassProbabilityCI |
(*args, **kwargs) |
A pass probability with the error bar its sample size actually earns. |
block_length_sensitivity |
(result: 'BacktestResult', *, params: 'CombineParams', lengths: 'Sequence[int]' = (1, 5, 10, 20), paths: 'int' = 1000, horizon_days: 'int \| None' = None, seed: 'int' = 0) -> 'BlockLengthSensitivity' |
Re-run the Monte Carlo across block lengths and report the swing. |
BlockLengthSensitivity |
(*args, **kwargs) |
The same estimate at several block lengths, plus how far it moved. |
monte_carlo_by_year |
(result: 'BacktestResult', *, params: 'CombineParams', paths: 'int' = 1000, horizon_days: 'int \| None' = None, block_length: 'int' = 5, seed: 'int' = 0) -> 'YearStratification' |
One Monte Carlo per calendar year of the source run. |
YearStratification |
(*args, **kwargs) |
Per-year estimates, ascending by year. |
crosscheck |
(mc: 'MonteCarloResult', sweep: 'WindowSweep') -> 'CrossCheck' |
Compare a Monte Carlo against a window sweep, outcome by outcome. |
CrossCheck |
(*args, **kwargs) |
The bootstrap and the window sweep, forced to answer side by side. |
Sequential Combines
| Name | Signature | What it does |
|---|---|---|
sequential_combines |
(bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'WindowSweep' |
Run one fresh Combine per non-overlapping window_days-day window. |
spaced_combines |
(bars: 'Sequence[Bar]', factory: 'Callable[[], Strategy]', *, window_days: 'int', periods: 'int', account: 'AccountSize' = <AccountSize.S50K: '50K'>, dll_enabled: 'bool' = False, warm_start: 'bool' = True, validate: 'bool' = True) -> 'SpacedSweep' |
Run periods fresh Combines with start days spread evenly over the tape. |
WindowSweep |
(*args, **kwargs) |
Every window's attempt, plus the rates over them. |
SpacedSweep |
(*args, **kwargs) |
A requested number of periods, start days spread evenly, overlap allowed. |
WindowResult |
(*args, **kwargs) |
One window's Combine attempt, resolved on its own merits. |
Overfitting guards
| Name | Signature | What it does |
|---|---|---|
deflated_sharpe |
(daily_pnl: 'Sequence[Decimal]', *, trials: 'int', trial_sharpe_variance: 'float \| None' = None) -> 'DeflatedSharpe \| None' |
Deflate a strategy's Sharpe by the number of trials it was selected from. |
DeflatedSharpe |
(*args, **kwargs) |
A Sharpe ratio and what it is worth once the search is accounted for. |
TrialLedger |
(path: 'str \| Path') |
A file-backed count of every strategy configuration you have tried. |
probability_of_backtest_overfitting |
(matrix: 'Sequence[Sequence[float]]', *, splits: 'int' = 16) -> 'PBOResult' |
Run CSCV over a trials x observations performance matrix. |
PBOResult |
(*args, **kwargs) |
Probability of Backtest Overfitting, by combinatorially symmetric cross-validation (Bailey, Borwein, Lopez de… |
sharpe_ratio |
(daily_pnl: 'Sequence[Decimal]') -> 'float \| None' |
Per-day Sharpe of a daily P&L series. NOT annualized. |
Search & walk-forward
| Name | Signature | What it does |
|---|---|---|
optimize |
(bars: 'Sequence[Bar]', factory: 'StrategyFactory', grid: 'Sequence[Mapping[str, object]]', *, account: 'AccountSize' = <AccountSize.S50K: '50K'>, objective: 'Callable[[Report], Decimal]' = net_pnl_objective) -> 'OptimizationResult' |
Run every configuration in grid over bars and keep them all. |
OptimizationResult |
(*args, **kwargs) |
Every configuration tried, plus the winner — in that order of emphasis. |
walk_forward |
(bars: 'Sequence[Bar]', factory: 'StrategyFactory', grid: 'Sequence[Mapping[str, object]]', *, folds: 'int' = 4, account: 'AccountSize' = <AccountSize.S50K: '50K'>, objective: 'Callable[[Report], Decimal]' = net_pnl_objective) -> 'WalkForwardResult' |
Anchored walk-forward: select on everything before a fold, score on it. |
WalkForwardResult |
(*args, **kwargs) |
Anchored walk-forward: each fold re-selects on all data before it. |
FoldResult |
(*args, **kwargs) |
One walk-forward step: chosen in-sample, scored out-of-sample. |
net_pnl_objective |
(report: 'Report') -> 'Decimal' |
Default objective: net P&L after all fees. |
Attempt economics
| Name | Signature | What it does |
|---|---|---|
evaluate_ev |
(mc: 'MonteCarloResult', *, economics: 'EvalEconomics') -> 'EvalEV' |
Combine a bootstrapped outcome distribution with your own prices. |
EvalEV |
(*args, **kwargs) |
Expected value of one attempt, and the number that needs no assumption. |
EvalEconomics |
(*args, **kwargs) |
What one evaluation attempt costs YOU. No defaults on the prices. |
Continuous contracts
| Name | Signature | What it does |
|---|---|---|
stitch_continuous |
(series: 'Mapping[str, Sequence[Bar]]', *, spec: 'InstrumentSpec', symbol: 'str \| None' = None, roll_days: 'Mapping[str, date] \| None' = None) -> 'ContinuousSeries' |
Back-adjust several expiries into one continuous series. |
ContinuousSeries |
(*args, **kwargs) |
A stitched, back-adjusted series ready to hand to Backtest. |
RollEvent |
(*args, **kwargs) |
One seam: the day the front month changed, and what it cost to align. |
SymbolStrategy methods (what you call inside on_bar)
| Name | Signature | What it does |
|---|---|---|
SymbolStrategy.use |
(self, indicator: 'T', *, session: 'Session \| None' = None) -> 'T' |
Register an indicator: auto-updated on every matching bar and |
SymbolStrategy.buy |
(self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None' |
Buy this contract (market unless a price kwarg implies otherwise); |
SymbolStrategy.sell |
(self, size: 'int', *, stop_loss_ticks: 'int \| None' = None, take_profit_ticks: 'int \| None' = None, limit_price: 'Decimal \| None' = None, stop_price: 'Decimal \| None' = None, custom_tag: 'str \| None' = None) -> 'int \| None' |
Sell this contract (market unless a price kwarg implies otherwise); |
SymbolStrategy.close |
(self) -> 'None' |
Flatten this contract's position; a rejection goes to on_reject. |
SymbolStrategy.cancel_working |
(self) -> 'None' |
Cancel every working order on this contract, one cancel per order; |
SymbolStrategy.move_stop |
(self, *, price: 'Decimal \| None' = None, ticks: 'int \| None' = None) -> 'int' |
Move every bracket stop on this contract; returns how many moved. |
SymbolStrategy.move_target |
(self, *, price: 'Decimal \| None' = None, ticks: 'int \| None' = None) -> 'int' |
Move every bracket take-profit on this contract; returns how many moved. |
SymbolStrategy.note |
(self, text: 'str') -> 'None' |
Attach a free-text breadcrumb to the current bar's replay frame. |
SymbolStrategy.on_bar |
(self, bar: 'Bar') -> 'None' |
— |
SymbolStrategy.on_fill |
(self, trade: 'HalfTradeModel') -> 'None' |
— |
SymbolStrategy.on_reject |
(self, error: 'APIError') -> 'None' |
Called with the APIError when a sugar order call is rejected. |
Indicators — every one wraps TA-Lib; none is reimplemented here.
| Name | Signature |
|---|---|
Adx |
(period: 'int' = 14, history: 'int \| None' = None) |
Atr |
(period: 'int', history: 'int \| None' = None) |
BBands |
(period: 'int' = 20, deviations: 'float' = 2.0, history: 'int \| None' = None) |
Cross |
(a: 'ValueSource', b: 'ValueSource') |
Ema |
(period: 'int', history: 'int \| None' = None) |
Highest |
(period: 'int', history: 'int \| None' = None) |
Indicator |
(*args, **kwargs) |
Lowest |
(period: 'int', history: 'int \| None' = None) |
Macd |
(fast: 'int' = 12, slow: 'int' = 26, signal: 'int' = 9, history: 'int \| None' = None) |
NotReadyError |
(*args, **kwargs) |
Obv |
(history: 'int \| None' = None) |
Rsi |
(period: 'int', history: 'int \| None' = None) |
Sma |
(period: 'int', history: 'int \| None' = None) |
StdDev |
(period: 'int', history: 'int \| None' = None) |
Stoch |
(fastk: 'int' = 5, slowk: 'int' = 3, slowd: 'int' = 3, history: 'int \| None' = None) |
TalibIndicator |
(name: 'str', price: 'str \| None' = None, history: 'int \| None' = None, **params: 'int \| float') |
TalibLine |
(owner: 'TalibIndicator', name: 'str') |
ValueSource |
(*args, **kwargs) |
WarmIndicator |
(*args, **kwargs) |
BacktestResult fields — the frozen run outcome.
| Field | Type |
|---|---|
verdict |
Verdict |
reason |
str |
ending_balance |
Decimal |
starting_balance |
Decimal |
profit_target |
Decimal |
floor |
Decimal |
best_day |
Decimal |
total_profit |
Decimal |
days_traded |
int |
day_records |
tuple[DayRecord, ...] |
breach |
Breach | None |
trade_count |
int |
equity_curve |
tuple[tuple[int, Decimal], ...] |
rejections |
tuple[tuple[int, int], ...] |
bar_equity |
tuple[BarEquity, ...] |
round_trips |
tuple[RoundTrip, ...] |
SummaryStats fields — derived metrics. See the basis notes below.
| Field | Type |
|---|---|
verdict |
Verdict |
closed_trades |
int |
win_rate |
Decimal | None |
expectancy |
Decimal | None |
profit_factor |
Decimal | None |
max_drawdown |
Decimal |
equity_peak |
Decimal |
final_balance |
Decimal |
net_pnl |
Decimal |
distance_to_floor |
Decimal |
consistency_headroom |
Decimal |
days_traded |
int |
exposure |
Decimal | None |
start_ts_ns |
int | None |
end_ts_ns |
int | None |
provisional |
bool |
avg_win |
Decimal | None |
avg_loss |
Decimal | None |
payoff_ratio |
Decimal | None |
expectancy_r |
Decimal | None |
longest_losing_streak |
int |
breakeven_cost_per_half_turn |
Decimal | None |
sortino |
Decimal | None |
calmar |
Decimal | None |
daily |
DailyStats |
drawdown |
DrawdownStats |
round_trips |
RoundTripStats |
4. The ctx seam¶
A strategy touches the venue through self.ctx and nothing else — never import a concrete
broker, clock or fill model. That is what makes the class run live unchanged.
ctx.orders |
place / buy / sell / modify / cancel / cancel_all / search_open / get / wait_for_fill — the exact SDK OrderResource signatures, incl. stop_loss_ticks/take_profit_ticks magnitudes and PlaceOrderBracket. Async; raise APIError on refusal. |
ctx.positions |
search_open / close / partial_close / close_all |
ctx.history |
retrieve_bars(cid, *, unit, unit_number, start_time, end_time, limit=1000, live=False, include_partial_bar=False) — serves only already-seen bars; non-native bar specs refused. |
ctx.clock |
now_ns() / now(). Never datetime.now(). |
ctx.account_id |
first positional argument to every orders/positions call |
ctx.instrument(cid) |
InstrumentSpec — tick size, tick value, session metadata |
SymbolStrategy is sugar over exactly that: buy/sell/close/cancel_working, plus
move_stop/move_target over the bracket children (stop_orders/target_orders are the
filtered views), self.position (flat/is_long/is_short/avg_price),
self.working_orders, self.spec, self.bars_gated. The sugar
does not raise — it catches APIError, routes it to on_reject and returns None;
self.ctx.orders is the raising path. Both are counted in the rejection tally (§5.5).
Hooks, all optional: on_start (sync), on_bar, on_order, on_fill, on_position,
on_stop (sync). Override on_*; drivers call handle_*, so framework bookkeeping cannot be
severed. The sim emits no PositionModel event on a FULL close (live sends a size-0 snapshot) —
detect flatness from self.position.flat in on_fill, not by overriding on_position.
5. Things that silently produce WRONG answers¶
Every item here is a bug that runs clean and returns a number.
5.1 Execution¶
- Market orders fill at the NEXT bar's open,
positions.close/partial_closeincluded. An order participates in a bar only ifaccepted_ts <= bar.ts_event(its OPEN). That is the no-look-ahead firewall, not latency modelling, and it is not configurable. wait_for_fillraises in sim. React inon_order/on_fill— parity-safe in both worlds.- Limit orders need the bar to trade THROUGH the level, ≥1 tick past. An exact touch,
including at the open, does not fill unless
BarFillConfig(fill_limit_on_touch=True)— a sensitivity probe, not a fix for "it should have filled". Stops gap-fill at the open when the bar opens beyond them, else trigger ±stop_slippage_ticks(default 1) adverse. - A bar holding both your stop and your target resolves adverse-extreme-first: the stop wins. One pessimistic O→H→L→C path orders everything by TRIGGER level, never by slippage-adjusted price; breach ties beat fills; equity is re-checked at each fill as it applies.
- Bracket children are created at the entry FILL (offsets from the ACTUAL fill price) and go active only the bar AFTER it — one bar where the entry is on and its protection is not. Reduce-only orders clamp to the live position and can never flip exposure.
OrderModel.trail_priceis the trail DISTANCE as a price offset, not the stop level.STOP_LIMITandJOIN_BID/JOIN_ASKare rejected at Tier-0.APIError.error_code: 4 position cap / account dead / day-locked, 5 no-trade window 16:10–18:00 ET or weekend, 8 unknown contract, 2 validation. They happen live too.
5.2 Indicators¶
Rules here; tables, per-function warmups and the refused-function list in docs/INDICATORS.md.
- Never hand-write an indicator. Every value comes from TA-Lib, so no formula can drift. 152
of its 161 functions are reachable via
TalibIndicator("NAME", …); the typed wrappers (Sma,Ema,Rsi,Atr,Macd,BBands,Stoch,Adx,Obv,StdDev,Highest,Lowest) are the same machinery with names. Nine functions are REFUSED at construction rather than failing silently mid-run (EXP/COSH/SINH/ACOS/ASIN,MAVP,MAXINDEX/MININDEX/MINMAXINDEX); so are lossy parameters (Sma(14.7),matype=2.9, bools) and per-function-invalid periods (Rsi(1)raises at construction, not 500 bars in). use()it or it never updates, and registration order IS update order — inputs before theCrossreading them. ACrossthat was neveruse()d RAISES fromup/downonce both inputs are ready; it used to readFalseforever and take zero trades in silence.- Register the OWNER, not a
TalibLine.macd.line("macdsignal")has noupdateon purpose anduse()on one raises;use()theMacd, thenuse(Cross(macd.line("macd"), macd.line("macdsignal"))). ACrossover aCrossis refused (ValueError), a non-Indicatorat registration (TypeError) rather than hours in. readyis NOTwarm.ready= a value exists (atlookback) and is theuse()gate.warm= the bounded buffer is FULL (athistory_bars), so the value no longer depends on where this run started. Parity begins atwarm: preloadhistory_barsbars live, notlookback. Never write "parity holds once ready". Readhistory_barsoff the instance — usuallymax(512, 64 × lookback), butSAR()derives it fromacceleration(→ 2000 on a lookback of 2), MAMA fromslowlimit, KAMA has a 9000-bar floor.- Values are
Decimalbut deliberately NOT tick-snapped — an indicator level is not a tradeable price. Never route one into grid math without an explicitround_to_tick. They are float64-precise, not Decimal-exact: deterministic across reruns, but aCrosson a Bollinger edge can flip onSTDDEV's cancellation noise. A band touch is not exact. use(ind, session=…)scopes the DATA;trade_sessions=scopes the DECISION. Two independent switches — see §5.11. Scoping an indicator does not restrict trading, and restricting trading does not starve an indicator.- One indicator instance per thread; a parameter sweep gets one per worker.
5.3 Data¶
stamphas no default and never will. It declares what your source timestamp MEANS:stamp="open"→ts_event = ts,ts_init = ts + step;stamp="close"→ts_init = ts,ts_event = ts - step. Getting it wrong is a silent one-bar look-ahead that inflates every result and raises nothing. Confirm your vendor's convention before you pass it.- Naive timestamps, off-grid prices and DAY-unit wrangling (the 23h Globex day needs a
session-aware resampler that does not exist) are rejected with the fix named.
validate=Trueis the default: any ERROR finding raisesDataValidationErrorcarryingissues, and turning it off does not disable the feed's time-order assert. - There is NO exchange holiday calendar, by decision — a hand-maintained one was wrong on
~five dates a year and its cleaner ran by DEFAULT, silently deleting tradable sessions. So a
holiday bar is indistinguishable from any weekday bar, the broker rejects weekends only, and
synthetic_bars(days=N)counts WEEKDAYS, so a tape spanning Thanksgiving or Good Friday emits a session real data would not contain. Filter exchange holidays upstream, in the data you feed in.
5.4 Rules — and the caveat that outranks the rest¶
- The MLL breach number is FRAMEWORK-COMPUTED and cannot be reconciled against the gateway.
The check runs continuously on realized plus unrealized equity, marked at the bar's adverse
extreme (open long at the low, open short at the high) so an intrabar breach is never missed.
The real SDK
PositionModelhas nounrealized_pnlandTradingAccountModelhas no equity or open-P&L field — there is no live number to diff it against. The value deciding pass/fail is computed here and nowhere else. Report every verdict as a diagnostic, never as an authoritative pass/fail; the rule and fee constants are cited config, uncalibrated against a real account (docs/topstep-rules.md§9). This is the largest parity hazard in the product. - MLL is two-state: the floor ratchets ONLY on end-of-day closed balance and locks permanently at the starting balance once EOD ≥ start + buffer, while breach checks run continuously. Intraday-trailing is Apex, not Topstep — do not port that intuition. A session close at or below the floor is itself a breach; flatten fees alone can do it.
- The DLL (
dll_enabled=True, off by default) is NOT a fail: flatten plus a day-lock cleared at 18:00 ET, no re-emission while locked, and it does not preclude passing. - Consistency:
best_day <= 0.5 * total_profit. Formula and per-size dollar table are both canonical; no cross-size ratio holds. Negativereport.stats.consistency_headroommeans a passing balance still fails the Combine. - The 16:10 ET flatten is forced liquidation at
forced_liq_slippage_ticksadverse plus $10 per contract. Close your own positions first if you care about the price.
5.5 A zero-trade result with a REJECTED line is a wiring bug¶
Refused placements are tallied, never silent: every order path funnels through one choke point
counting rejections by error_code into SimBroker.rejections and BacktestResult.rejections
(sorted (error_code, count) pairs), and print(report) emits a REJECTED line whenever it is
non-empty. A strategy whose every order was refused — cap 4, outside-hours 5 — used to render as
a clean zero-trade report. Always read that line: zero trades WITH rejections is broken
wiring; zero trades without is a quiet strategy. Same for report.bars_gated — if it equals your
bar count, your warmup is longer than your data.
5.6 Metrics carry a BASIS, and mixing bases is how a report lies¶
Every figure in SummaryStats states its basis in its docstring because most of them admit two
honest answers. The ones that bite:
- Gross vs net.
win_rate,expectancy,expectancy_r,payoff_ratio,avg_win,avg_lossandprofit_factorare GROSS;net_pnlis net of every fee. Fees are charged on every half-turn, so a strategy can printprofit_factor 1.14and still lose money — the bundledsma_crossexample does exactly that. Never quote a gross ratio besidenet_pnlas though they share a basis. - Half-turns, not round trips.
closed_tradescounts CLOSING HALF-TURNS. A flip is one half-turn that closes and reopens.breakeven_cost_per_half_turnis per half-turn for the same reason — that is how fees actually land. There is no round-trip pairing: it needs a FIFO convention whose live-gateway equivalence is unverified (docs/topstep-rules.md§9 box 5). - Two R-multiples, and they are not the same number.
stats.round_trips.expectancy_ris the TRUE one: net P&L over the dollars actually risked at entry (bracketstop_loss_ticksx tick value x size), per flat-to-flat excursion.stats.expectancy_ris a gross-basis approximation using R = average loss, kept for strategies that set no stops. Prefer the round-trip one whenround_trips.with_known_risk > 0; the gap between that andround_trips.countis how much of the R statistics to distrust. Risk is captured AT ENTRY — trailing the stop later does not change it. round_tripsare NET of fees; the half-turn figures are GROSS. The basis flips between the two blocks deliberately, because a round trip is a complete decision and the honest question is what it earned after costs. The bundledsma_crossexample prints grossprofit_factor 1.14and net round-tripexpectancy -3.81on the same run. The net one is the one that pays you. Counts differ too: a three-clip scale-out is threeclosed_tradesbut one round trip.round_trips.worst_rbelow -1 means a stop was jumped — a gap, or the Tier-0 model's stop slippage. Read it before trusting any risk-based sizing built on "I only ever lose 1R".round_trips.worst_tradeis the survival question;worst_ris the strategy question. R says how an excursion went against its own plan, dollars say whether the ACCOUNT could absorb it. A single trip that loses more than the Daily Loss Limit ends the day whatever its R looked like. Readworst_tradeagainst the DLL and againstdistance_to_floorbefore either ratio.round_trips.avg_duration_ns/max_duration_nsare nanoseconds, and both endpoints are fill stamps — Tier-0 stamps a fill at its bar's OPEN or CLOSE and never between, so an excursion opened and closed inside one bar reports0, not its true sub-bar life. Holding time is a rule question here: the engine flattens at 16:10 ET, so a typical duration approaching the session's remainder means the flatten, not the strategy, is closing the trades.- Three drawdowns, deliberately.
drawdown.static(fixed initial balance),drawdown.eod_trailing(Topstep's actual MLL mechanic) anddrawdown.intraday_trailing(Apex-style, ratchets on unrealized highs) are different numbers on the same path and a strategy can survive one while violating another. The legacymax_drawdownis a fourth, close-sampled measure. Say which one you mean. drawdown.avg_eod_trailingaverages EPISODES, not days — each excursion below the high-water mark measured at its own trough, overeod_episodesof them, witheod_trailingas the max of the same set. The gap between mean and max is the shape of the risk: a max far above the mean is one bad week, a max close to it means the curve lives at that depth and the worst case is the normal case.equity_peakis close-basis and seeded at the starting balance, so it never reports below it andequity_peak - max_drawdownis the trough that produced it. It matters because a trailing MLL floor is ANCHORED to a peak — but the real one ratchets on end-of-day closed balances, so the floor actually in force followed the EOD peak, which is at or below this. Where the distinction bites, quotedrawdown.eod_trailing.exposureis a FRACTION of bars, not a percentage, and it is what tells you how to read every other figure: identical drawdowns at 0.05 and 0.95 exposure are not the same risk. A bar counts when a trip was open strictly inside it, boundary resolved forward (a position opened at a bar's close belongs to the next bar). Multi-instrument runs count a timestamp ONCE — time with risk on, not a sum of per-symbol exposures.Nonewithout bars, which is not "never exposed". The denominator includes bars outside tradable hours.sortinoandcalmarare NOT annualized — daily-dollar and combine-horizon bases respectively. Annualizing a twenty-day sample produces a number with no defensible meaning.daily.stdevis dollars and population-basis (divides byclosed_days), matching sortino's downside deviation and likewise not annualized. Dollars because the DLL is a fixed dollar threshold, so this is the dispersion figure that makes "could I breach it" answerable. It is exactly0for a one-day run — an artifact ofn=1, not a finding.start_ts_ns/end_ts_ns/duration_nsare CALENDAR time, spanning every night and weekend the market was shut. Judge a run's length bydays_traded; the two diverge sharply on any sparse or gapped feed.provisionalisTruebelowPROVISIONAL_TRADE_FLOOR(200) closes and every rate, ratio and percentile is then an estimate swamped by sampling error. Nothing is suppressed — it is labelled, andprint(report)says so loudly. Do not quote a provisional metric as a finding.drawdown.intraday_trailingandmin_floor_headroomneedBacktestResult.bar_equity, and areNonewithout it (hand-built results, older runs). They are sampled over the modelled Tier-0 intrabar path, not real ticks — conservative estimates of a tick-based figure, not reproductions of one.min_floor_headroomis the closest the account ever came to termination;distance_to_flooris only where it ENDED.
5.7 The Monte-Carlo answers a different question, and has its own traps¶
metrics.monte_carlo(result, params=..., paths=..., seed=...) block-bootstraps the run's
observed trading days and replays each synthetic sequence through a fresh CombineKernel.
- The kernel decides every verdict. Nothing in that module re-implements the MLL ratchet, the consistency test or the pass condition — it drives the same three calls the live engine uses. Never add a second rulebook there.
- Read the autopsy, not the pass probability.
mll_breach/consistency_blocked/target_not_reachedpartition the failures and imply different fixes: resize, throttle the outsized day, or accept the edge is too slow. The headline number alone tells you nothing actionable. block_length=1is a footgun. It degenerates to an i.i.d. resample, destroys the losing streaks that actually blow accounts, and will report a pass probability that is far too kind. Default is 5.horizon_daysdefaults toBILLING_MONTH_DAYS(21), never the observed day count. A Combine has no time limit, only a monthly fee, so "one attempt" defaults to one fee cycle — the same unitEvalEconomicsbills in. The fixed default is also the safety property: a blown run stops recording days at the breach, so its count is a survival time, not a Combine length, and is never used as a horizon.source_truncatedstill flags that such a sample is survivorship-biased by construction — the days after the blow-up do not exist, every figure is conditioned on having survived, and nothing can repair that; do not quote such a result without the caveat.provisionalbelow 30 source days. Resampling cannot create information.- The point estimate ships with its own cross-examination (
metrics/confidence.py).pass_probability_cidouble-bootstraps the source days themselves — the error bar the day count earns, which the path count never was.block_length_sensitivityshows whether the streak assumption is load-bearing.monte_carlo_by_yearrefuses to average a hostile year against a kind one.crosscheckcompares againstsequential_combinesunder a binomial 2×SE null — when they disagree, the disagreement is the finding. Quote a pass probability with its CI, not alone;mc_confidencebundles the lot andreport.to_html(path, confidence=..., crosscheck=...)renders the cards. - It cannot invent a regime the tape never contained, and it inherits every uncalibrated constant (§5.4). A probability to three decimals from unverified inputs is precise, not accurate.
5.8 One long backtest answers a question nobody is asked¶
Nobody runs a single continuous evaluation for three years. They attempt a Combine, and if it
resolves, they attempt another. metrics/windows.py::sequential_combines models that: the tape
is cut into consecutive non-overlapping windows of window_days trading days, each run as a
completely fresh evaluation (new balance, new floor, new broker, new strategy instance).
- It is the empirical twin of the Monte-Carlo, biased the opposite way.
monte_carloresamples observed days — many paths, real sequencing destroyed beyond the block length. This preserves sequencing and regime exactly and pays in sample size: three years is ~35 independent 21-day windows, not 35,000. They shareclassify_failureso the autopsies are directly comparable. Run both; disagreement is the finding. - Non-overlapping is not a limitation, it is the point. Stepping one day at a time would give ~700 windows from three years, each sharing 20 of 21 days with its neighbour — the precision of 700 observations carrying the information of 35. The step is fixed at the window length rather than shipped as a footgun with a warning.
pass_rateis an estimate off a handful of samples. At ~35 attempts its standard error is near 8 percentage points before anything else is considered. Always read it besideattempts.- Warm starts happen OUTSIDE the engine, and must. Preload bars fed through a
Backtestwould land inday_recordsas flat days and moveclosed_days, every daily percentile,stdev,sortinoand the drawdown durations while leaving P&L untouched. SoSymbolStrategy.prewarmdrives the preceding history through the indicators only — noon_bar, no orders — and the run then sees only its own window.prewarmraises if the strategy is already bound. history_bars, notwarmup, is the preload size (§5.2'sreadyvswarm). That is real money:Sma(30)is warm atmax(512, 64 x 30) = 1920bars, which on RTH-only 5-minute candles is ~25 trading days of history per window. Expect the first windows of any tape to befully_warm=False— cold windows UNDER-trade, so leaving them in biasespass_rateDOWN.trailing_days_droppedis reported, never absorbed. A partial window is not an attempt, and scoring one would count a short evaluation as a failure to reach the target.
5.9 Optimizing is a search, and a search lies by default¶
optimize() was absent from this project on purpose — maximizing over a Combine metric is
an overfitting machine. It exists now because the guards do. Use it accordingly:
OptimizationResult.bestis not evidence. Selecting a maximum from a search guarantees a flattering number. The API keeps every trial specifically so you can feedpbo_matrixtoprobability_of_backtest_overfittingand the winner'sdaily_pnltodeflated_sharpe. A bare best-params answer is the thing to refuse.- PBO and walk-forward are not substitutes. PBO asks whether your SELECTION generalises and is symmetric in time, so it says nothing about decay. Walk-forward asks whether the edge SURVIVES FORWARD and is directional, so it conflates a decaying edge with a bad selection rule. Run both; they disagree in informative ways.
- Read
efficiency, not the OOS number. How much of the in-sample edge survived is the question. Negative efficiency — chosen for making money, then lost money — is the signature of a fit to noise. It isNoneagainst a non-positive in-sample objective, because a ratio against a loss inverts the sign of good and bad. consistent_foldsguards the aggregate. A good total built from one huge fold and three losers is not a robust strategy.- The default objective is net P&L, deliberately not the verdict. Optimizing a binary pass/fail throws away nearly all the information in a run and rewards configurations that scraped over the line once.
- Cost: a grid of G over F folds runs
G x (F + 1)complete backtests. Serial by design — a search that silently sampled would be worse than a slow one.
5.10 Multi-year runs need stitching, and the seam corrupts INDICATORS not P&L¶
A two-year run spans ~8 quarterly rolls. data/continuous.py::stitch_continuous turns
per-expiry bars into one back-adjusted series labelled with the bare ticker ("MNQ", which
spec_for_symbol already resolves), so the engine sees one instrument and one unbroken
account.
- A raw splice does NOT produce wrong P&L. The 16:10 flatten plus the session-roll
backstop mean no position and no working order survives a day boundary, and a roll seam IS
a day boundary — so entry and exit always share one contract. What the seam corrupts is
indicator state, which does span it: a spurious
Cross, anAtrspike inflated for a whole lookback, a falseHighest/Lowestbreakout. The resulting trades are priced correctly and should never have been taken. - Additive back-adjustment is therefore exactly P&L-neutral, because entry and exit
share the offset and it cancels in the difference. Pinned end to end by
test_engine_pnl_is_invariant_under_a_constant_price_shift— if that ever fails, additive stitching is unsafe and this reasoning is wrong. - Never ratio-adjust. Multiplicative offsets do not cancel (P&L would move) and push
prices off the tick grid, which
SimBrokerasserts on every fill. - What adjustment does distort: logic keyed to ABSOLUTE price levels (round numbers, a
fixed price target). Tick-relative logic —
stop_loss_ticks,take_profit_ticks, all the strategy sugar — is unaffected. Adjusted prices also no longer match livehistory.retrieve_bars, which returns raw single-expiry quotes. - The roll spread is measured at ONE shared instant from both expiries, so the two contracts must overlap; a spread taken across a time gap folds that interval's market move into the adjustment permanently. Rolls are forced onto day boundaries — that is what makes the cancellation argument hold.
- Still your problem on a multi-year tape: exchange holidays (~20/year, no calendar ships — filter upstream) and the uncalibrated constants (§5.4).
5.11 Sessions scope indicator DATA and trading DECISIONS, separately¶
ASIA / LONDON / NEW_YORK live in core/sessions.py; the worked example is
examples/session_scoped.py.
- The two switches are independent, and conflating them is the bug.
use(Atr(14), session=NEW_YORK)restricts which bars that indicator is computed from.SymbolStrategy(..., trade_sessions=(NEW_YORK,))restricts whenon_barmay fire. Indicators advance regardless oftrade_sessions— an indicator fed only the tradable window develops gaps and computes a different value from the same tape. Both default to today's behaviour, and an unscoped strategy is byte-identical to one written before sessions existed. - Which indicators to scope is a modelling decision, and it is not uniform. Dispersion
measures (
Atr,StdDev,Rsi,Stoch,BBands) describe how much price moves PER BAR, and that is session-dependent: on a 24h feed anAtr(14)read at 09:30 ET is computed almost entirely from thin pre-market bars, so it understates NY volatility exactly when stop distance is being sized. Level measures (Sma,Ema) answer where price IS, and the overnight move is real — an NY-onlyEmais anchored to yesterday's 16:00 close. Levels are continuous across sessions; dispersion is not. - A scoped indicator warms in ITS OWN cadence. It needs
history_barsbars of its session, so on 5-minute bars an NY-scopedSma(30)spans ~25 trading days against ~7 unscoped. Readstrategy.warm(per-indicator update counts) orhistory_bars_by_session— neverbars_seen >= history_bars, which cannot express two cadences and OVERSTATES warmth.sequential_combinesstill SIZES its preload slice from the unscopedhistory_bars, so a scoped strategy will honestly reportfully_warm=Falserather than silently lying. - A
Crossinherits its inputs' scope and refuses a conflictingsession=. Mixed-scope inputs are refused outright: they advance on different bars, so comparing them compares values sampled at unrelated instants. - Sessions are defined in their own timezone, not as fixed ET offsets. DST comes from the
IANA database, so nothing rots — London and New York switch on different dates, and the London
window really is 04:00 ET rather than 03:00 for ~3 weeks each spring and ~1 each autumn. A
fixed ET block is still one line:
Session("LONDON_ET", ET, time(3), time(11)). - Session membership is a DIFFERENT axis from
trading_day_of(). Asia sits after the 18:00 ET rollover, so its bars belong to the NEXT trading day. Membership is tested onts_event(the bar's OPEN) over a half-open[start, end)window, so the 09:29→09:30 bar — every print of it pre-market — is not New York. trade_sessionsonly narrows what THIS strategy does. It never widens what the venue permits: the 16:10 ET flatten and the 16:10–18:00 no-trade window apply either way.- You need 24h data. The shipped
data/sample_mnq_1m.csvis RTH-only (390 bars/day, 09:30–15:59 ET) and contains no Asia or London bars, so nothing session-scoped is observable on it. For synthetic bars passsynthetic_bars(..., hours="globex"); the default"rth"mode is the New York session and nothing else. - Not built: per-session performance attribution. Nothing in
SummaryStatssplits P&L by session, so scoping is currently a modelling choice you make, not one the report scores.
6. The end-to-end workflow¶
What using this framework actually looks like, in the order you do it:
from topstep_backtest import AccountSize, Backtest
from topstep_backtest.metrics import crosscheck, mc_confidence, sequential_combines
from topstep_backtest.rules.params import combine_params
# 1. Write the strategy (§2), then run ONE backtest — recorded, so the replay
# tab can show intent against execution.
report = Backtest(bars, MyStrategy(CONTRACT), account=AccountSize.S50K, record=True).run()
print(report) # verdict, day trail, four statistics blocks
# 2. Sanity-check the wiring BEFORE reading any number.
assert not report.result.rejections # zero trades + rejections = broken wiring (§5.5)
assert report.bars_gated != len(bars) # warmup longer than your data
# 3. Read the run, minding the basis (§5.6) — and step the replay tab to each
# entry: the bracket where you meant it, the notes matching the fills.
s = report.stats
if s.provisional:
... # under 200 closes: estimates, not findings
s.round_trips.expectancy # NET per trade — the one that pays you
s.round_trips.expectancy_r # true R, when with_known_risk > 0
s.drawdown.eod_trailing # Topstep's actual MLL mechanic
s.drawdown.min_floor_headroom # closest the account came to death
s.daily.p05 # the bad day to size against
# 4. Stop trusting one sample — real attempts first, resampled ones second.
sweep = sequential_combines(bars, lambda: MyStrategy(CONTRACT), window_days=21)
c = mc_confidence(report.result, params=combine_params(AccountSize.S50K), seed=7)
check = crosscheck(c.mc, sweep) # same billing-month horizon on both sides (§5.7)
# 5. Act on the AUTOPSY, and quote the CI, never the bare point (§5.7).
# mll_breach -> resize
# consistency_blocked -> throttle the outsized day; the edge is fine
# target_not_reached -> the edge is too slow; nothing risk-side helps
c.ci.p05, c.ci.p95 # the error bar the source-day count earns
c.sensitivity.spread # wide = the streak assumption is doing the work
check.divergent # bootstrap vs real windows: disagreement is the finding
report.to_html("run.html", confidence=c, crosscheck=check) # the cards, archived
# 6. Was it an edge, or did you search until something looked good? The trial
# count must be RECORDED — a remembered one is always too low, because the
# sweeps you abandoned are exactly the ones that inflated the winner.
ledger = TrialLedger(".trials.json")
ledger.record("sma-10-30", sharpe=sharpe_ratio(daily))
dsr = deflated_sharpe(daily, trials=ledger.count(), trial_sharpe_variance=ledger.sharpe_variance())
dsr.survives # deflated >= 0.95, the conventional bar
dsr.expected_max_sharpe > dsr.sharpe # the search alone explains the result
# 7. Is the attempt worth its price? Every figure is yours; none is baked in.
ev = evaluate_ev(
c.mc,
economics=EvalEconomics(
monthly_fee=Decimal("149"), pass_value=Decimal("2000"), reset_fee=Decimal("99")
),
)
ev.breakeven_pass_value # lead with THIS — it assumes nothing
Steps 2 and 3 are not optional politeness. A zero-trade run with rejections and a gross-basis ratio quoted as if it were net are the two most common ways to read a confidently wrong result out of this framework.
Steps 6 and 7 have one trap each. DSR with trials=1 deflates nothing — reporting it is
how a user reveals they never counted the search, and the statistic exists precisely to stop
that. EV's pass_value is an assumption you supply and this package cannot check; it
dominates the answer, so quote breakeven_pass_value ("a pass must be worth at least $X")
rather than ev unless you can defend the input.
Runnable end to end: examples/run_combine.py (steps 1–3), examples/run_windows.py and
examples/run_montecarlo.py (steps 4–5, the latter at two position sizes so the autopsy
visibly discriminates). docs/TUTORIAL_EMA_CROSSOVER.md walks all of it line by line, and
website/workflow.md is the narrative version with the gate each stage must pass.
7. What Backtest() refuses, and what to use instead¶
Coming from backtesting.py, the dialect is
deliberately close but causal. Backtest(bars, s, cash=50_000, commission=0.002) raises
ValueError: Backtest() does not support '<knob>': … naming the alternative. Verbatim, from
harness.py::_REJECTED_KNOBS:
| Knob | The error's explanation |
|---|---|
cash= |
account economics are fixed by the Combine — pick account=AccountSize.S50K/S100K/S150K instead |
commission= |
fees are the real Topstep schedule (fills.fees.TopstepFees), applied always; they are not a per-trade knob. To correct a rate against your own blotter, pass the whole schedule: fee_model=TopstepFees(overrides={...}) |
margin= |
there is no margin model — the Combine's MLL floor and position cap (combine_params) are the risk limits |
spread= |
execution costs come from the Tier-0 bar fill model (fills.bar_fill.BarFillConfig slippage ticks), not a synthetic spread |
trade_on_close= |
filling at the decided bar's close is the exact look-ahead the accepted_ts firewall forbids; orders fill from the NEXT bar |
hedging= |
the venue nets positions per contract — hedged positions cannot exist |
exclusive_orders= |
hidden auto-close orders would mutate the strategy's intent sequence (the parity gate); make reversals explicit awaits |
finalize_trades= |
end-of-session positions are governed by the 16:10 ET flatten rule, never a stats toggle |
Also refused: the strategy class instead of an instance (TypeError naming the fix);
fractional or equity-fraction sizing (size is int contracts under a hard cap); init()-time
full-array precompute (self.I), since there is no full array at the live edge; and synchronous
order placement — the SDK is async and the awaits ARE the live contract. Translations:
| backtesting.py | here |
|---|---|
init() / next() |
__init__ / async def on_bar(self, bar) |
self.I(SMA, self.data.Close, n) |
self.fast = self.use(Sma(n)) |
| arrays lose their NaNs | the ready-gate; count in report.bars_gated, pin with warmup= |
self.data.Close[-1] |
the bar argument; ctx.history.retrieve_bars(…) for seen history |
crossover(fast, slow) |
self.use(Cross(fast, slow)) → .up / .down |
self.buy(size=2, sl=…, tp=…) |
await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80) |
trade.sl = price / trade.tp = price |
await self.move_stop(price=…) / move_target(price=…) — or ticks= from position.avg_price, signed in the position's favour. These amend REAL working orders, so a rejection is real too |
stats = bt.run() → 30-key Series |
report = bt.run() → Report + .stats/.result/.trades |
optimize() |
metrics.optimize(...) — but read §5.9 first; it keeps every trial on purpose |
plot() |
report.show() — interactive HTML tearsheet (report.to_html(path) to name the file; bt.run_with_tearsheet(dir) to get one from every run; Backtest(..., record=True) adds the bar-by-bar replay tab) |
8. Knobs — the calibration seam¶
Four keyword arguments, and only these, move economics or pessimism.
Backtest(
bars,
strategy,
account=AccountSize.S100K, # 50K/100K/150K — balance, floor, target, position cap
dll_enabled=True, # optional Daily Loss Limit (off by default)
fill_config=BarFillConfig( # topstep_backtest.fills.bar_fill
stop_slippage_ticks=1, # adverse ticks on EVERY triggered stop (gap or intrabar)
market_slippage_ticks=0, # adverse ticks on market fills
fill_limit_on_touch=False, # True = optimistic touch fills (a sensitivity probe)
),
broker_config=SimBrokerConfig( # topstep_backtest.execution.sim_broker
forced_liq_slippage_ticks=2, # MLL/DLL/16:10 market-dump slippage
liquidation_fee_per_contract=Decimal("10"), # Topstep's auto-liquidation fee
max_trail_ticks=1000,
history_depth=20000,
),
fee_model=TopstepFees(
overrides={
"MNQ": FeeSchedule( # topstep_backtest.fills.fees
commission_per_side=Decimal("0.25"),
exchange_per_side=Decimal("0.37"),
nfa_per_side=Decimal("0.02"),
)
}
),
)
fee_model= is a calibration seam, not a P&L dial: the shipped rates are researched but
uncalibrated, and this is how you substitute what your own blotter charges. It is the only fee
input; per-trade commission= is refused (§7). Every default sits on the conservative side —
moving one toward optimism is a probe to report next to the baseline, not a new baseline.
9. Checklist before returning a strategy¶
- Subclasses
SymbolStrategy/Strategy; reaches the venue only viaself.ctxor the sugar; no concrete broker, nodatetime.now(). - Every indicator is
self.use()d — inputs before theirCross, owners not lines. on_bar/on_order/on_fillareasync def; every order call isawaited.- Entry guarded on
self.position.flat and not self.working_orders. - Size is an
intnumber of contracts, within the account cap. - Ran end to end once and you read the report — verdict, REJECTED line,
bars_gated. report.result.rejectionsis empty, or every code in it is one you intended.- On real data:
stamp=set from the vendor's documented convention, holidays filtered upstream. - The verdict is reported as a diagnostic, not a pass/fail (§5.4).
10. Map: module → what it owns¶
| Module | Owns |
|---|---|
protocols.py |
Clock / Bar / DataFeed / Broker / OrderApi / PositionApi / HistoryApi / FillModel / FeeModel / Fill / WorkingOrder |
core/ |
money (tick grid, P&L), instruments (15-product spec table), time (ET sessions), ids |
rules/ |
CombineKernel — THE rulebook state machine, golden-fixtured; params per account size |
clock/ |
deterministic TestClock, wall-clock LiveClock |
fills/ |
Tier-0 BarFillModel/BarFillConfig, shared build_path, TopstepFees/FeeSchedule |
data/ |
wrangler (bars_from_dataframe/bars_from_records, explicit stamp), validator, synthetic, feed (ListBarFeed), clean (infer_symbol_from_filename/AmbiguousSymbolError — filename→symbol inference when loading CSVs) |
execution/ |
SimBroker/SimBrokerConfig (lifecycle, OCO brackets, netting, breach walk, forced liq), rejections |
engine/ |
BacktestEngine loop, BacktestResult (msgspec-serializable, carries rejections) |
strategy/ |
Strategy + StrategyContext (the write-once seam), SymbolStrategy (use(), ready-gate, sugar), tracker (NetPosition/PositionTracker/OrderTracker, pure folds over SDK models) |
metrics/stats.py |
SummaryStats + compute_summary — every derived figure off one run; DrawdownStats, DailyStats, RoundTripStats. Bases in §5.6 |
metrics/montecarlo.py |
monte_carlo — pass probability + violation autopsy, replayed through the REAL CombineKernel, never a second rulebook (§5.7) |
metrics/overfitting.py |
deflated_sharpe, TrialLedger, probability_of_backtest_overfitting (§5.9). Float, not Decimal — normal-theory statistics |
metrics/windows.py |
sequential_combines — one long tape replayed as consecutive independent Combine attempts, warm-started off the preceding history (§5.8) |
metrics/walkforward.py |
optimize (keeps EVERY trial) + anchored walk_forward + split_by_day (§5.9) |
metrics/economics.py |
evaluate_ev + EvalEconomics — EV per attempt from YOUR prices; lead with breakeven_pass_value |
data/continuous.py |
stitch_continuous — many expiries → one back-adjusted series on the bare ticker (§5.10) |
indicators/ |
TA-Lib adapter (152 of 161 functions), typed wrappers, Cross, Indicator/ValueSource |
replay.py |
Recorder + Replay — bar-by-bar run recording behind Backtest(record=True): decisions at the ctx seam (via recording proxies), events, indicator values, state tracks, sparse SummaryStats snapshots built by the SAME metrics/stats helpers (terminal snapshot pinned equal to compute_summary). Observation only: results are byte-identical either way |
harness.py |
Backtest facade (one shared clock, instruments from the feed, strict validation, refuses backtesting.py knobs) and Report |
11. Not built — do not write code against it¶
docs/ROADMAP.md is the authority. These names appear in older notes and other frameworks and
do not exist here: LiveBroker, RecordingLiveBroker, DataEngine, TrailingMaxLossLimit,
prob_fill_on_limit (the real field is BarFillConfig.fill_limit_on_touch). Also absent: a
Parquet/Arrow data catalog, multi-timeframe resampling, a MessageBus, per-year / per-regime /
per-session breakdowns (§5.11), higher fill tiers, an XFA rule set, and any holiday calendar.
Do write code against these — they used to be on the list above and now ship: Monte-Carlo
(§5.7), optimize() and the overfitting guards (§5.9), the sequential-Combine sweep (§5.8),
the continuous-contract stitcher (§5.10), and the HTML tearsheet (report.show() /
report.to_html(path) / Backtest.run_with_tearsheet(dir), which writes a timestamped file
per run; print(report) remains the text render). Their signatures are in the generated
table in §3.
Parity status, precisely: structural conformance is proven — a pyright-strict protocol
conformance test plus a signature diff of SimOrderApi.place against the SDK's
OrderResource.place (tests/parity/test_broker_conformance.py), both run in CI. The
behavioural gate — an intent-sequence test asserting an identical Submit/Modify/Cancel
sequence sim vs live — is Phase 4 and is not met. Behavioural parity is argued, not proven.
12. Modifying the framework?¶
Read docs/DESIGN.md first: it carries the architecture contract this file assumes — Decimal
on the tick grid with FIFO position accounting, int-ns UTC time with ET boundaries, the
accepted_ts no-look-ahead firewall, intrabar determinism along one pessimistic path, and the
rule that protocols.py is THE canonical interface module (change a signature there first or
not at all).
The gate is seven commands, and all seven run in CI (.github/workflows/ci.yml, on Python
3.12/3.13/3.14, plus a job that installs the built wheel into a clean venv and smoke-tests it
and a job that builds the documentation site):
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv run python scripts/gen_api_surface.py --check # §3 is stale
uv run python scripts/check_doc_links.py # a doc link dangles
uv run mkdocs build --strict # needs --extra docs
ruff is pinned >=0.16,<0.17 on purpose — format output changes between minors, so a floating
pin turns ruff format --check red on an unrelated upgrade. pyproject.toml declares exactly
three extras, [data], [dev] and [docs]. If you change the public surface, regenerate §3
with uv run python scripts/gen_api_surface.py — and add the new names to GROUPS in that
script, or they will be documented nowhere and the --check will still pass.
The documentation site¶
uv run --extra docs mkdocs serve renders it locally at http://127.0.0.1:8000. mkdocs.yml
sets docs_dir: website, which holds ONLY the site-original pages (index, quickstart,
results, beyond-one-backtest). Everything else is assembled at build time by
scripts/gen_docs.py:
- The repo's own markdown (
README.md, this file,CHANGELOG.md,docs/*.md) and the runnable examples are mirrored into the build at their repo-relative paths. That layout is load-bearing —docs/DESIGN.mdlinks to../AGENTS.mdand the tutorial links to../examples/ema_cross.py, and mirroring keeps ONE link graph that works in a checkout, in the sdist and on the site. Never "fix" a link for the site alone; you would fork the graph and rot whichever copy nobody reads. - The API reference is one mkdocstrings stub per module, listed in
MODULESin that script. A new public module needs an entry there and anav:entry inmkdocs.yml; miss the first and it is documented nowhere, miss the second and the strict build fails.
show_if_no_docstring: true is REQUIRED in the mkdocstrings options, not cosmetic: this codebase
shares one docstring across groups of related fields (p05..p95, best_r/worst_r), griffe
attaches it to the LAST name only, and the default would silently drop every other member.