Roadmap¶
This file answers one question: is X built yet? Rule numbers cited in the table are defined
in topstep-rules.md.
| Phase | Deliverable | Status |
|---|---|---|
| 0. Accounting spine & protocols | core/money.py (Decimal↔int-tick), core/instruments.py, frozen protocols.py |
Done |
| 1. Rule kernel | Two-state MLL + optional DLL + consistency + position cap, golden-fixtured | Done — rules/kernel.py::CombineKernel |
| 2. Engine skeleton | TestClock, single ts_init queue, three-phase settle, no-look-ahead |
Done — the non-decreasing-ts_init assert lives on engine/backtest.py::BacktestEngine |
| 3. SimBroker + Tier-0 fills + rules wired | fills/bar_fill.py::BarFillModel, fills/fees.py::TopstepFees, netting/PnL, intrabar mark path, forced liquidation |
Done — the mark/breach walk and forced liquidation are inside SimBroker, not separate components |
| 4. Strategy API + parity gate | Strategy/StrategyContext, SymbolStrategy, reference MA-cross strategies |
Partial — API and examples/{sma,ema}_cross.py ship, plus session scoping (core/sessions.py: per-indicator use(session=) for input data and trade_sessions= for decisions, two independent switches; examples/session_scoped.py); the intent-sequence gate is NOT met. Only structural parity is proven (pyright-strict protocol conformance + a place() keyword diff, tests/parity/test_broker_conformance.py) |
| 5. Data layer | Parquet/Arrow catalog, validator, warmup, causal continuous-contract stitcher | Partial — data/validator.py, SymbolStrategy warmup and data/continuous.py::stitch_continuous ship (additive back-adjustment, volume or explicit rolls, tick-exact offsets, RollEvent metadata). No Parquet/Arrow catalog |
| 6. Analytics | Decimal-path metrics + tearsheet; Monte-Carlo pass-probability; overfitting guards (PBO/DSR/walk-forward) | Partial — metrics/stats.py::compute_summary covers trade stats (expectancy/payoff/streaks/breakeven cost), flat-to-flat round trips with true R-multiples plus dollar extremes and holding times, drawdown under all three prop conventions plus duration/recovery/mean-episode-depth/min-floor-headroom, the daily-P&L distribution with a dollar standard deviation, exposure, the equity peak, the run window, and Sortino/Calmar. metrics/montecarlo.py::monte_carlo block-bootstraps observed days through the real CombineKernel for pass probability, a violation autopsy (MLL breach / consistency-blocked / target-not-reached) and days-to-target. metrics/overfitting.py::deflated_sharpe + TrialLedger deflate a Sharpe by the recorded trial count, and metrics/economics.py::evaluate_ev turns a pass probability plus YOUR prices into an EV and a breakeven. metrics/walkforward.py adds optimize (which keeps every trial, not just the winner), anchored walk_forward with efficiency, and metrics/overfitting.py::probability_of_backtest_overfitting (CSCV). optimize() was deferred for a long time as an overfitting machine; it ships now because DSR and PBO exist to catch what it produces. The interactive HTML tearsheet ships: report.to_html(path) / report.show() render one self-contained file (candlestick tape with fills marked, equity vs the trailing MLL floor, daily P&L, R-multiple distribution, and every text-render stat with its basis label; tearsheet/, charting via vendored Lightweight Charts). Bar-by-bar replay ships: Backtest(record=True) records every decision/event/indicator value/running-stats snapshot (replay.py, observation-only, golden-pinned byte-identical results) and the tearsheet grows a scrubber over it. metrics/confidence.py qualifies the Monte-Carlo number itself: pass_probability_ci (double bootstrap — the error bar the source-day count earns), block_length_sensitivity (does the estimate depend on the one arbitrary knob), monte_carlo_by_year (per-year strata so a hostile year is never averaged against a kind one — the per-year breakdown this row long listed as the remainder), and crosscheck (bootstrap vs sequential_combines, binomial 2×SE under the MC null; when they disagree, the disagreement is the finding). mc_confidence bundles the first three and report.to_html(path, confidence=..., crosscheck=...) renders the cards. Remaining: per-regime (non-calendar) strata — per-YEAR ships; a volatility-tercile classifier was deliberately not invented |
| 7. Live adapter + calibration | LiveBroker shim over topstep-sdk, captured-gateway fixtures, field-level parity diff, calibration vs a real eval account |
Not started — nothing has run against a real account; fee and rule constants are uncalibrated |
| 8. Higher fill tiers | QuoteFillModel → DepthFillModel → MBOFillModel (CME FIFO) |
Not started — Tier-0 only |
| — Parked: Express Funded (XFA) | Funded phase (scaling, payout paths, post-first-payout MLL→0) as a new RuleSet |
Parked — only after the Combine path is trustworthy |
Not built — do not write code against these¶
RecordingLiveBroker,LiveBroker— no live or recording broker exists.SimBrokeris the onlyBroker.DataEngine— no such class. The ordering assert is onBacktestEngine.TrailingMaxLossLimit— no such class. The MLL lives insideCombineKernel.prob_fill_on_limit— no such field. The real knob isBarFillConfig.fill_limit_on_touch.- Parquet/Arrow data catalog, per-regime (non-calendar) and per-session breakdowns
(sessions scope indicator data and trading, but no metric splits P&L by session),
QuoteFillModel/DepthFillModel/MBOFillModel, XFARuleSet.
Built since this list was first written — the entries below are gone from it on purpose:
the continuous-contract stitcher (data/continuous.py::stitch_continuous), Monte-Carlo
(metrics/montecarlo.py::monte_carlo) with its confidence instruments and per-year strata
(metrics/confidence.py), the overfitting guards (metrics/overfitting.py,
metrics/walkforward.py), and the HTML tearsheet (tearsheet/, via report.to_html(path)
or report.show()). Write code against them.
Dropped: exchange calendar. There is no holiday/early-close table and none is planned; a
hand-maintained one shipped once, was wrong on roughly five dates a year, and silently
dropped tradable half sessions. Weekends and the daily 17:00–18:00 ET maintenance halt are
still modeled; full closures must be filtered upstream (rationale in AGENTS.md and the
CHANGELOG).