Changelog¶
All notable changes to this project are documented here. This project adheres to Semantic Versioning. While the version is below 1.0, minor releases may contain breaking changes.
0.4.0 — 2026-08-25¶
A minor: new features throughout. One behavioural default changed — the Monte-Carlo horizon now defaults to one billing month rather than the observed day count.
Added¶
- Regional sessions (
core/sessions.py:Session,ASIA,LONDON,NEW_YORK) and the two independent switches that use them.use(indicator, session=...)scopes an indicator's INPUT DATA;SymbolStrategy(trade_sessions=...)scopes DECISIONS. Keeping them separate is the whole feature — a strategy can hold a continuous 24hEmabeside an NY-onlyAtrand still trade only New York. Indicators advance regardless oftrade_sessions, because starving one outside the tradable window would leave it with gaps and a different value than the same indicator on the same tape. Both default to today's behaviour and an unscoped run is byte-identical.
Why you would scope at all: dispersion measures (Atr, StdDev, Rsi) describe how much
price moves per bar, and that is session-dependent — on 24h bars an Atr(14) read at 09:30
ET is computed almost entirely from thin pre-market bars, understating NY volatility exactly
when stop distance is being sized. Level measures (Sma, Ema) answer where price is, and
the overnight move is real: an NY-only Ema is anchored to yesterday's 16:00 close and blind
to a London repricing. Levels are continuous across sessions; dispersion is not.
Sessions are defined in their own local timezone, not as fixed ET offsets, so DST comes
from the IANA database and there is no hand-maintained table to rot — the same reasoning that
dropped the exchange calendar. London and New York switch on different dates, so the London
window really is 04:00 ET rather than 03:00 for about three weeks a year, and the tests assert
that against concrete 2026 dates. A fixed-ET block remains one line away
(Session("LONDON_ET", ET, time(3), time(11))), midnight-wrapping included.
Session membership is a different axis from trading_day_of(): Asia sits after the 18:00
ET rollover, so its bars belong to the next trading day. Membership is tested on ts_event
(the bar's open) over a half-open window, so the 09:29->09:30 bar — every print of it
pre-market — is not New York.
-
SymbolStrategy.warm,.registrations,.history_bars_by_session,.bars_out_of_session,.bars_seen,.in_trade_session()— the diagnostics scoping makes necessary. A scoped indicator warms in its OWN cadence, so it needshistory_barsbars of its session: on 5-minute bars an NY-scopedSma(30)spans ~25 trading days against ~7 for the same indicator on a 24h feed.warmcounts each indicator's own updates, which a single bar count cannot express, andsequential_combinesnow asks the strategy rather than comparing counts — the oldpreload >= neededsilently OVERSTATED warmth for a scoped indicator. -
synthetic_bars(hours="globex")— the full 23-hour electronic session (18:00 ET previous calendar day to 17:00 ET), reusing the canonical open fromcore/time.py. The default"rth"mode is unchanged. Without this nothing session-scoped is testable: the shippeddata/sample_mnq_1m.csvis RTH-only (exactly 390 bars/day, 09:30-15:59 ET) and contains no Asia or London bars at all. -
A
Crossnow inherits its inputs' data scope and refuses a conflicting one, alongside the existing refusal of unregistered inputs. Crossing a 24hEmaagainst an NY-scopedEmacompares values sampled on unrelated cadences; leaving theCrosscontinuous over scoped inputs would have it re-read unchanged values on most bars. Both were silent before. -
metrics.spaced_combines— pick how many Combine attempts to simulate. Wheresequential_combineslets the tape dictate the attempt count (disjoint windows), the new sweep takesperiodsandwindow_daysand spreads that many start days evenly across the tape, overlapping as much as the arithmetic requires — 100 days, 10 periods of 40 days starts an attempt roughly every week. Each period is still a completely fresh Combine (same shared implementation: day-aligned slicing, prewarmed indicators, the kernel's own verdict,classify_failureattribution), so the sweep measures how much passing depends on when the attempt starts. Because overlapping periods are not independent samples, the result carrieseffective_independent_windows,stride_daysandoverlap_fractionalongside its rates, and asking for more periods than the tape has distinct start days is refused rather than replaying identical windows as new observations.examples/run_spaced.pyrenders a sweep as a text report — the rates with their effective-sample caveat in the header, then one line per period in start order so start-date sensitivity is visible at a glance. -
An HTML tearsheet for sweeps.
sweep.to_html(path)/sweep.show()on bothWindowSweepandSpacedSweep(ortearsheet.render_sweep_html) render one self-contained page: a calendar timeline with each attempt drawn over its actual dates and stacked into lanes where they overlap — the reason overlapping attempts are not independent samples made visible rather than footnoted — the outcome autopsy as a stacked bar with each failure mode's prescription, every attempt's cumulative P&L overlaid from $0, a closest-approach-to-the-MLL-floor strip (a pass with $40 of headroom looks identical to a robust one in the pass rate; not here), and a sortable per-attempt table with P&L sparklines, all hover-linked. To feed the charts,WindowResultnow carriesdaily_pnl(and acumulative_pnlproperty), the same per-day series walk-forward's trials keep. No charting library and no embedded JSON: the charts are SVG generated in Python, so the render stays a pure byte-identical function of the frozen sweep data, and the evidence caveat is stamped in the page header where a screenshot cannot shed it. -
Confidence instruments for the Monte-Carlo pass probability (
metrics/confidence.py). The point estimate's simulation error was never the real uncertainty; these quantify what is.pass_probability_cidouble-bootstraps the observed day set itself and reports a 5th–95th percentile band — the error bar the source-day count earns (a 65% from 500 days and a 65% from 40 now read differently).block_length_sensitivityre-runs the estimate across block lengths and reports the spread: wide means the streak-clustering assumption is doing the work.monte_carlo_by_yearstratifies by calendar year so a hostile year is never averaged against a kind one — the spread across strata is the error bar non-stationarity imposes.crosscheckforces the bootstrap andsequential_combines(opposite biases, sharedclassify_failureon purpose) to answer side by side, flagging any outcome where the real windows sit outside 2×SE of what the bootstrap's own probability would produce; when they disagree, the disagreement is the finding.mc_confidencebundles the first three, and every estimate runs through the samemonte_carlo_from_blockscore as the number it qualifies. All deterministic per seed, like everything else inmetrics/. - Monte-Carlo cards on the tearsheet.
report.to_html(path, confidence=..., crosscheck=...)(andshow/to_timestamped_html) render the estimate with its CI, the block-sensitivity row, the per-year strata, and the bootstrap-vs-windows card. Caller-computed on purpose: writing a file never triggers thousands of simulations as a side effect, and the render stays a pure, byte-identical function of its inputs. -
The workflow, documented (
website/workflow.md, on the site nav as "The workflow"). The stages in the order the questions become answerable — envelope, data, strategy, one recorded run, real attempts, the distribution with its confidence, the search guards, the price — each ending with the gate that must pass before the next stage's number means anything.AGENTS.md§6 is the code-first version and now walks the same loop (recorded run,sequential_combines,mc_confidence+crosscheck, guards, EV). -
A native loader for databento-data-playground Parquet exports.
topstep_backtest.data.loaders.load_bars(path)reads an export produced by that project'sconvert.pyand returns(bars, spec, meta)ready forBacktest(...). The product, bar span and timestamp stamping are read from the file's embedded Parquet metadata rather than passed in — each is a silent corruption when guessed — and the file's copy of the instrument economics is cross-checked against the engine'sInstrumentSpec, failing loudly on any mismatch instead of picking one. Files without the metadata blob are refused. Requires the existing[data]extra (pandas + pyarrow); nothing new to install.examples/run_real_data.pyauto-detects such exports and loads them with no declaration flags (and refuses the flags on one — re-declaring what the file states is the contradiction they exist to prevent), and the module joins the site's API reference and quickstart.
Changed¶
- The Monte-Carlo horizon defaults to one billing month (
BILLING_MONTH_DAYS = 21), not the observed day count. A Combine has no time limit, only a monthly fee, so "one attempt" defaults to one fee cycle — the same unitEvalEconomicsbills in (itstrading_days_per_monthdefault now is this constant). The fixed default also retires the guard that madehorizon_daysmandatory for a blown source run: the old hazard was defaulting to a survival time, and the new default cannot. A failed run still reportssource_truncated, and the survivorship-bias caveat stands unchanged. sample_day_path,nearest_rank, andmonte_carlo_from_blocksare module-level names inmetrics/montecarlo.py(previously underscore-private) so the confidence instruments run through the identical core; they remain outside__all__.- The replay cockpit is easier to read. The running-stats panel no longer stacks all
~50 rows: snapshots are regrouped into by-type cards — status, trades (gross),
round trips (net), risk & quant, daily P&L — shown one at a time behind filter
chips, with a
•on any chip whose hidden card the newest snapshot moved. The values are the same rows the results tab renders, relocated rather than rebuilt (a stat missing from the regrouping map fails loudly instead of silently vanishing from the replay). The event log moved from the cramped right column to a full-width strip under the charts — one event per line, heading and kind filters on one row — and the panel text was tightened throughout. Both tabs' price charts also gained a marker legend (entry/exit arrows, and in the replay the cross fires and the working stop/target/avg-entry price lines), so the glyphs on the tape no longer require guessing.
0.3.0 — 2026-08-20¶
A minor: new features, no intended breaking changes.
Added¶
- Bar-by-bar replay of a run.
Backtest(..., record=True)hooks a recorder into the engine andReport.replaycarries the full recording (replay.py): every decision at thectxseam with its exact parameters and the broker's answer — rejections keep their gateway code and message instead of being flattened to a count — every order event, fill (gross P&L + costs), position snapshot, session enforcement (16:10 flatten, session close, MLL breach, DLL lockout), per-bar indicator values named by the attribute the strategy stores them under (self.fast = self.use(Sma(20))records asfast; multi-output indicators record line by line;Crossfires are events), warmup-gated bars as a state rather than a hole, and change-only tracks of position / working orders / balance folded from the same SDK events the strategy's own views fold. Recording is observation only: results are byte-identical with it on or off, andStrategy.note()— the new narrative breadcrumb channel (self.note("why")inside any hook, tagged with the hook that said it) — is a pure sink that cannot influence a run. Both pinned by goldens (tests/golden/test_replay_goldens.py). - Running statistics that cannot drift from the report. The recording carries sparse
SummaryStatssnapshots (one per closing trade, day close, and breach — never per bar), each computed over the run's prefix by the samemetrics.statscodecompute_summaryuses: the shared helpers were extracted for this (compute_trade_close_stats,compute_daily_stats,compute_round_trip_stats,compute_eod_drawdown,sortino_ratio,exposure_fraction—compute_summary's behavior is unchanged), and the only hand-rolled parts are the O(1) equity folds. The terminal snapshot is golden-pinned byte-equal tocompute_summaryand property-swept across tapes (tests/property/test_replay_props.py). One documented exception:exposurerefreshes at day closes only (the single prefix stat with no exact incremental form) and is exact again at the terminal snapshot. - The tearsheet grows a replay tab when the report carries a recording. The page splits
into results (the finished run) and replay · bar by bar, a cockpit sized to the
viewport — charts down one side, the settled state and running stats beside them, the
event log alongside — so stepping a run never means scrolling between the tape and the
panel that explains it. Step it bar by bar (buttons, slider, ←/→ with shift for ×10,
Home/End, space to autoplay at 1, 5, 15 or 40 bars a second, click a chart or a log row
to jump); everything after the
cursor is veiled on every one of the replay tab's charts, so it shows only what the
strategy had seen, while the results tab keeps showing the run that finished. A strip
across the tape carries the live trade numbers at the cursor — position and average entry,
open P&L, the day, equity, floor headroom, and the closed-trade record so far (net P&L,
net win rate, trades) lifted from the running-stats snapshot in force, plus the cursor
bar's OHLC; the toolbar's
statsbutton hides it. Charts follow the cursor in a selectable window (whole run, or 60–480 bars — a long recording opens following, since fitting 30,000 bars into a pane gives each one half a pixel), and their series' last-value badges are off in favour of a price line at the cursor's own close and equity: a badge reporting the end of a run the cursor has not reached is future information printed on the axis of the one view that promises not to show any. The running stats mark every row the newest snapshot moved (hover for the previous value) and say which snapshot of how many is in force; the event log filters by kind, with counts, and says how many rows are hidden. Panels show the settled state at the cursor (position, working orders — also drawn as price lines on the candlestick pane — balance, equity, floor headroom, today's P&L), the indicator lines on their own synced chart with cross fires marked on the tape, the running-stats snapshot in force, and a filterable-by-eye event log with rejections loud. The open tab rides in the URL fragment (#replay), so a reload — or a link to an archived file — comes back to the view it was left on. All figures are Python-preformatted by the same helpers as the rest of the page; the JS computes nothing beyond differences of two displayed numbers. Long tapes embed a loudly labelled window pastREPLAY_AUTO_FRAME_LIMIT(20,000) frames — breach-centred when there is a breach, else the tail — andto_html(path, replay=(start, end) | "full" | "off")overrides it (show(),to_timestamped_html()andrun_with_tearsheet()forward the same knob). Payload schema version bumps to 2. - Amending a live bracket, and the average entry price to measure it from.
SymbolStrategy.move_stop(...)/move_target(...)amend the reduce-only stop and limit the venue creates when an entry fills, taking eitherprice=(absolute, passed to the broker as given) orticks=— an offset from the average entry signed in the position's favour, soticks=0is breakeven andticks=10is ten ticks of locked profit long or short. They act on bracket children only (parent_order_idset), so a stop-entry of your own is never mistaken for protection; they move every child (a scale-in has several) and return the count; rejections go toon_rejectas with every other sugar call. A level measured off the average is snapped to the tick grid against the position, because the venue's average is quantized to tick/100 and a stop rounding toward profit would lock in a tick the entry cannot support.stop_orders/target_ordersexpose the same filtered view. Amended levels are live from the next bar, which theaccepted_tsfirewall guarantees and an end-to-end test pins by exit price. position.avg_price— the venue's display average (PositionModel.average_price), which the position view previously discarded. Carried from snapshots rather than re-derived, so no second averaging convention is invented for a number the gateway publishes; fills keep it honest where they can do so exactly (opening from flat, flipping through it, flattening), and a scale-in leaves the previous average standing until the snapshot corrects it — the next event in the sim, one hub round trip live.Report.replay_json(path)dumps the raw recording as JSON — for diffing two runs or verifying what was captured without a browser in the way.SymbolStrategygains a read-onlyregistered_indicators;SimBrokergains O(1)last_bar_equity.examples/run_replay.pyshows the whole flow, narrated.
[0.2.3] — 2026-08-16¶
Fixed¶
- The
[data]and[dev]extras now installpyarrow, so Parquet input works out of the box.examples/run_real_data.pyhas always branched topd.read_parqueton a.parquet/.pqsuffix, but pandas ships no Parquet engine of its own — so[data]alone met that branch withImportError: Unable to find a usable engine, on the format most vendors (Databento included) actually export. Deliberately unbounded, unliketopstep-sdkandta-lib: a breakingpyarrowmajor fails loudly at read time instead of quietly changing a number. Reading the same bars from CSV and from Parquet produces a byte-identical report.
Released as 0.2.3, not 0.2.2: the v0.2.2 tag was cut from a commit three minutes before
this change merged, so 0.2.2 shipped without pyarrow and its Parquet path fails exactly
as 0.2.1's did. A PyPI version can never be re-uploaded, so the fix moved forward a version
rather than the tag moving backward.
[0.2.2] — 2026-08-16¶
Fixed¶
- Repository plumbing only; no user-facing change.
ruff formatalso formats Python inside markdown fences, so the tearsheet examples added to the README andwebsite/results.mdin 0.2.1 were unformatted code and CI had been red since — through two merges. And the docs workflow'sactions/configure-pagesstep called the Pages REST API under acontents: readtoken, failing every deploy with "Resource not accessible by integration" while blaming the repository's Pages settings; MkDocs takes its base URL fromsite_url, so the step was removed rather than the token widened.
[0.2.1] — 2026-08-16¶
Fixed¶
- The README claimed shipped subsystems do not exist. Its status section told readers
there is "no walk-forward, no PBO/DSR overfitting guard, and no EV-per-attempt model" —
all three have shipped since 0.2.0 (
metrics/walkforward.py,metrics/overfitting.py,metrics/economics.py), anddocs/ROADMAP.mdand the site documented them correctly the whole time. This is the rot the 0.2.0 doc pass fixed everywhere except the one file a new reader opens first, and it was the package's PyPI project description. The bullet now keeps the limitation that is still true — every analytic resamples the tape you supplied, so none of them invents a regime your data never contained — without the false premise. The "Next:" pointer drops the analytics that shipped and names what the roadmap actually lists as remaining. - The HTML tearsheet was missing from the README and from the site's "Reading the report"
page — the headline feature of 0.2.0, absent from both places a reader looks for it. Both
now cover it, including
Backtest.run_with_tearsheet(dir), which was documented only in the changelog andexamples/run_tearsheet.py.AGENTS.md§7 and §11 pick it up too, and the examples list in the README no longer omitsrun_tearsheet.py. - Version strings that describe the CURRENT state (README status,
docs/DESIGN.md's "proven today", the captured provenance lines in the tutorial andwebsite/results.md) said 0.1.0. The 0.1.0 references that are historical — when the exchange calendar was removed, and why — are left alone deliberately.
[0.2.0] — 2026-08-16¶
Added¶
- An interactive HTML tearsheet —
report.to_html(path)writes ONE self-contained document (no server, no CDN, opens offline): the candlestick tape with every fill marked per round trip, the equity curve against the trailing MLL floor with the modelled intrabar envelope, daily P&L, the R-multiple distribution, and every stats section the text render prints, label for label. The labels and money formatting come from the same helpers asReport.__str__(now shared in_render.py), so the two renders cannot disagree on a basis.report.show()is the no-path variant: it writes a freshtopstep-tearsheet-*.htmltemp file — the one documented write not handed an explicit path — and opens the default browser. Rendering is a pure function of the frozen report data (byte-identical across calls, no wall clock), and the page carries the provenance caveat and the!! PROVISIONALbanner exactly as the text render does. Charting by TradingView Lightweight Charts™ v5.2.1, vendored undertearsheet/_assets/with its Apache-2.0 license and a visible attribution in the footer; zero new Python dependencies.Reportnow also carriesbarsandinstruments(keyword-only, defaulted — the pre-tearsheet constructor signature still works) so the price panes can draw the tape.examples/run_tearsheet.pyshows the flow. Backtest.run_with_tearsheet(directory)andReport.to_timestamped_html(directory)— a sheet from EVERY run without naming a file each time. The wrapper is exactlyrun()followed by the writer, so nothing about the run changes; the writer names the filetearsheet-<UTC stamp>.html, creates missing directories, and suffixes-2,-3, ... rather than overwriting when two runs land in the same second. The wall clock is read for the FILENAME only — the document is still the byte-identical pure render of frozen report data, which a test pins against a plainrun().to_html().metrics.sequential_combines— replay a long tape as consecutive independent Combine attempts rather than one continuous multi-year run, which is what a trader actually does. Cuts the tape into non-overlapping windows of N trading days and runs each as a fresh evaluation (new balance, floor, broker and strategy), returning per-window verdicts plus pass / MLL-breach / consistency-blocked / target-not-reached rates. Sharesclassify_failurewithmonte_carloso the two autopsies are directly comparable — they estimate the same quantity by opposite methods, and disagreement between them is the finding. Non-overlapping only, deliberately: a one-day step would yield the precision of 700 observations carrying the information of about 35.SymbolStrategy.prewarm(bars)andSymbolStrategy.history_bars— drive history through the registered indicators without trading, so a window can start warm. This has to happen outside the engine: preload bars fed through aBacktestland inday_recordsas flat days and moveclosed_days, every daily percentile,stdev,sortinoand the drawdown durations while leaving P&L untouched.history_barsis the preload size (warm, not merelyready) —Sma(30)needs 1920 bars, which on RTH-only 5-minute candles is ~25 trading days per window.metrics.classify_failure— promoted from private so the sweep and the Monte-Carlo cannot drift apart on how a failure is attributed.indicators.WarmIndicator— the protocol for an indicator that declareshistory_bars. Runtime-checkable because aCrossholds no history of its own.- A documentation site (
mkdocs.yml+website/, MkDocs Material).uv run --extra docs mkdocs serveto read it locally. Three new hand-written pages — a quickstart, Reading the report (what every figure in the output means, its basis, and which ones mislead when read alone, including what this framework deliberately does not report and why), and Beyond one backtest — plus a full API reference generated from the live docstrings byscripts/gen_docs.py, so a signature on the site cannot drift from the one in the code. The existing markdown and the runnable examples are mirrored into the build at their repo-relative paths, so one link graph serves the checkout, the sdist and the site. Built with--strictin CI, which makes an unresolved link or a dangling#anchora build failure. SummaryStats.exposure— the fraction of the run's bars that held a position. Identical drawdowns at 0.05 and 0.95 exposure are not the same risk, and nothing else in the summary distinguished them. A bar counts when a round trip was open strictly inside it, boundary resolved forward; multi-instrument runs count a timestamp once (time with risk on, not a sum of per-symbol exposures).Nonewithout bars, which is not the same claim as "never exposed".SummaryStats.equity_peak— the close-basis high-water mark, seeded at the starting balance. A trailing MLL floor is anchored to a peak, so this is the number that set the floor the run had to stay above. The real ratchet is end-of-day, sodrawdown.eod_trailingremains the figure to quote where that distinction matters.SummaryStats.start_ts_ns/end_ts_nsand theduration_nsproperty — the run's window, so a serialized summary carries the span it describes. Calendar time, closes included; judge length bydays_traded.RoundTripStats.best_trade/worst_trade— the NET DOLLAR extremes, alongside the existing R extremes. R says how an excursion went against its own plan; dollars say whether the account could absorb it. Readworst_tradeagainst the DLL anddistance_to_floorfirst.RoundTripStats.avg_duration_ns/max_duration_ns— holding time per excursion. Tier-0 stamps a fill at its bar's open or close and never between, so a trip opened and closed inside one bar reports0.DrawdownStats.avg_eod_trailing/eod_episodes— mean depth over the run's separate excursions below the high-water mark, each at its own trough, witheod_trailingas the max of the same set. A max far above the mean is one bad week; a max close to it means the worst case is the normal case.DailyStats.stdev— daily P&L dispersion in dollars, population basis, not annualized. The DLL is a fixed dollar threshold, so this is what makes "could I breach it" answerable rather than a hope.SummaryStats.verdict(pluspassed/failedproperties) — the combine outcome now travels with the metrics it describes. Three-state:IN_PROGRESSmeans the tape ran out before the combine resolved and is the outcome of most runs, sopassedandfailedare BOTHFalsethere. Never readnot passedas "failed".- Trade statistics on
SummaryStats:avg_win,avg_loss,payoff_ratio,expectancy_r,longest_losing_streak,breakeven_cost_per_half_turn,sortino,calmar. Gross-basis where the existing ratios are gross-basis; Sortino and Calmar are deliberately not annualized.expectancy_ruses R = average loss, not per-trade initial risk (see AGENTS.md §5.6). SummaryStats.provisionalandPROVISIONAL_TRADE_FLOOR(200) — below the floor every rate, ratio and percentile is an estimate swamped by sampling error. Metrics are still computed;print(report)labels them loudly rather than letting a six-trade run read as a finding.DrawdownStats— drawdown under all three conventions prop firms use (staticfrom the initial balance,eod_trailingwhich is Topstep's actual MLL mechanic, andintraday_trailingwhich ratchets on unrealized highs), pluslongest_days,time_to_recovery_days, andmin_floor_headroom— the closest the account ever came to termination, as opposed todistance_to_floor, which is only where it ended.DailyStats— per-day P&L distribution (worst/best, win/lose/flat counts, p05/p25/median/ p75/p95 by nearest rank) andbest_day_pct_of_profit, the consistency-rule ratio.-
BacktestResult.bar_equity/SimBroker.bar_equity— per-bar realized+unrealized equity envelope and the MLL floor in force, which is what makes intraday-trailing drawdown and distance-to-floor-over-time computable. Sampled over the modelled Tier-0 intrabar path, not real ticks: a conservative estimate, not a reproduction. Defaults to empty, and dependent metrics reportNonerather than guessing. -
RoundTrip/BacktestResult.round_trips/SimBroker.round_trips— flat-to-flat excursions, the colloquial "per trade" basis as against the half-turns the gateway reports. Boundaries need no new convention: a trip opens at flat-to-positioned and closes at the return to flat, and a flip closes one and opens the next, with the flip's half-turn and costs attributed to the trip it CLOSED. This is a reporting grouping over the broker's existing FIFO half-turns — it never re-derives P&L, so the open gateway-pairing question (§9) cannot move these numbers without moving the half-turns first. - True R-multiples.
RoundTrip.initial_riskcaptures the dollars at risk at entry from the bracket stop (stop_loss_ticksx tick value x size), andr_multipleis net P&L over it.None— never 0 — when any opening fill carried no stop, because a signal-exit strategy has no defined risk and guessing one manufactures an R out of nothing. Captured AT ENTRY, so trailing the stop afterwards does not change R. -
RoundTripStatsonSummaryStats.round_trips— count,with_known_risk, and NET-basis win rate, avg win/loss, payoff ratio, expectancy,expectancy_r(the true R),best_r/worst_r.worst_rbelow -1 means a stop was jumped by a gap or slippage. -
metrics/montecarlo.py::monte_carlo— pass probability and violation autopsy by replaying block-bootstrapped trading days through the REALCombineKernel. Nothing re-implements the rulebook: each synthetic path drives the kernel through the same three calls the live engine uses, so a rule fix lands here for free and this can never drift from the engine. - Block bootstrap, not i.i.d. Trading days cluster, and a trailing drawdown is precisely a
bet against clustering; sampling days independently would break up the losing streaks that
blow accounts and report a pass probability far too kind.
block_lengthdefaults to 5. - Intraday excursion travels with each day (from
bar_equity), so an MLL breach on the way to a green close is reproduced when that day is replayed at a different balance. Resampling closing balances alone would miss every intraday breach. - The autopsy partitions every path into
pass/mll_breach/consistency_blocked/target_not_reached, which imply different fixes: resize, throttle day-level lumpiness, or accept the edge is too slow for the window. expected_days_to_pass/median_days_to_pass(conditional on passing), terminal-balance p05/median/p95, anddll_lock_rate. Deterministic givenseed; refuses an empty sample rather than reporting a confident 0% on no evidence; flagsprovisionalbelow 30 source days.- Refuses to default
horizon_dayswhen the source run FAILED. A blown run stops recording days at the breach, so that count is a survival time, not a Combine length — defaulting to it silently simulated one-day Combines for a strategy that died on day one.MonteCarloResult.source_truncatedadditionally flags that such a sample is survivorship-biased by construction: the days that would have followed the blow-up do not exist, so every figure is conditioned on having survived that long. metrics/overfitting.py— Deflated Sharpe Ratio (Bailey & Lopez de Prado): how impressive an observed Sharpe is given how many configurations it was selected from.expected_max_sharpeis what the best of N random strategies would show under the null — when the observed Sharpe is below it, the search alone explains the result.TrialLedgeris a file-backed, human-auditable count of every configuration tried, because a self-reported trial count is always too low: the abandoned sweeps are exactly the ones that inflate the winner. Deliberately float (normal-theory statistic over skew and kurtosis), per-period and NOT annualized, and no new dependency — stdlibNormalDist.metrics/economics.py— EV per evaluation attempt from aMonteCarloResultplusEvalEconomics. No price is baked in: fees and payouts change and are exactly the class of number this project refuses to hard-code uncalibrated. Subscriptions bill in whole months (a 22-day attempt is two), a failing attempt is billed for the full horizon, and reset fees are charged only on the paths that fail. The headline isbreakeven_pass_value— how much a pass must be worth to justify attempting — because it needs no assumption about payouts, unlikeev, whosepass_valueinput dominates it.breakeven_pass_probabilityis solved rather than divided, since expected cost is itself a function of the pass probability.data/continuous.py::stitch_continuous— continuous-contract stitching. Turns per-expiry bars into one back-adjusted series labelled with the bare product ticker ("MNQ"), whichspec_for_symbolalready resolved, so the engine sees one instrument and one unbroken account across a multi-year tape.- A raw splice does not corrupt P&L, contrary to the usual worry: the 16:10 flatten
plus the session-roll backstop mean no position or working order survives a day
boundary, and a roll seam is a day boundary — entry and exit always share one contract.
What a seam corrupts is INDICATOR state, which does span it (spurious
Cross,Atrspike inflated for a lookback, falseHighest/Lowestbreakout). - That is also why additive back-adjustment is exactly P&L-neutral: entry and exit
share the offset and it cancels in the difference. Pinned end to end by
test_engine_pnl_is_invariant_under_a_constant_price_shift, which asserts identical net P&L, verdict, MLL floor and per-day P&L under a constant price shift. - Additive only, never ratio: a multiplicative offset does not cancel and pushes prices
off the tick grid that
SimBrokerasserts on every fill. Offsets are whole tick counts by construction, since both expiries trade the same grid. - Rolls by volume (first overlap day the successor out-trades the incumbent, applied
monotonically so it cannot flap) or explicit via
roll_days, and always forced onto trading-day boundaries — a mid-session roll would break the cancellation argument. Spreads are measured at ONE shared instant from both expiries; a spread taken across a time gap would fold that interval's market move into the adjustment permanently. RollEventrecords every seam (day, both contracts, both prices, tick offset) andContinuousSeries.roll_daysexposes them, since an indicator's lookback is still crossing a seam forlookbackbars afterwards.metrics/walkforward.py—optimize()and anchored walk-forward.optimizewas deliberately absent for a long time (maximizing over a Combine metric is an overfitting machine); it ships now because DSR and PBO exist to catch what it produces, and it is built to hand you that evidence: it returns EVERY configuration's result plus apbo_matrix, never just a winner.walk_forwardnever scores a configuration on the data that selected it, and leads withefficiency(the share of the in-sample edge that survived) rather than the raw OOS number — negative efficiency is the signature of a fit to noise.efficiencyisNoneagainst a non-positive in-sample objective, because a ratio against a loss inverts the sign of good and bad. Aggregate efficiency sums before dividing so one tiny-denominator fold cannot dominate, andconsistent_foldscatches a good total built from one huge fold and three losers. Splits are day-aligned (split_by_day): cutting mid-session would score half a day as a day.metrics/overfitting.py::probability_of_backtest_overfitting— PBO via CSCV (Bailey, Borwein, Lopez de Prado & Zhu). Asks a sharper question than DSR: when you pick the best configuration in-sample, does it stay good out-of-sample, or were you picking noise? PBO and walk-forward are not substitutes — PBO is symmetric in time and says nothing about decay; walk-forward is directional and conflates decay with a bad selection rule. Block sums are precomputed so any union of blocks is scored exactly without re-walking the series.-
examples/run_montecarlo.py— runs the same SMA cross at 2 and at 30 contracts to show the autopsy discriminating: identical edge, but 100% "target not reached" at small size versus 100% "MLL breach" at large size. The size does not change the edge; it changes which way the strategy fails, and therefore what you would do about it. -
scripts/check_doc_links.py+ a CI step — every relative link and#anchoracross the shipped markdown must resolve. A dangling doc pointer is worse than no pointer: an agent that follows one concludes the guidance does not exist and invents its own. External URLs are not fetched, deliberately — a network call would make the check flaky and someone else's 404 is not a defect this repository can fix.
Changed¶
- The generated API table in
AGENTS.md§3 now covers the analytics surface and the stitcher (monte_carlo,deflated_sharpe,TrialLedger,probability_of_backtest_overfitting,optimize,walk_forward,evaluate_ev,stitch_continuousand their result types). They shipped documented only in hand-written prose, which is exactly the rot the generator exists to prevent — and--checkpassed the whole time, because it only ever verified the names it was told about. gen_api_surface.pynormalizes a function-valued default (objective=net_pnl_objective) to its bare name. It previously rendered as<function … at 0x…>, a per-process address that would have turned--checkred on source nobody touched — the same failure mode as the version string it already guards against. A regenerated table containing any surviving address now fails the script loudly.AGENTS.md§11 anddocs/ROADMAP.mdlisted the contract stitcher, Monte-Carlo and the overfitting guards under "not built — do not write code against these", while other sections of the same two files documented them with module paths. Corrected, and both lists now say explicitly which entries moved off them.docs/DESIGN.md§3.8 stated anti-overfitting as a requirement for a hypothetical future sweep. That sweep shipped; the section now names the code that enforces each clause.print(report)gaineddrawdown,round tripsanddaily P&Lsections, the provisional banner, and the new trade statistics.- Verdict golden artifacts regenerated twice, for
bar_equityand thenround_trips. Each time the new field was verified to be the only changed field in both artifacts — verdicts, balances and every other field byte-identical, so no behaviour moved.
0.1.0 — 2026-07-28¶
First public release. Pre-alpha: the engine core is well tested, but the Topstep
rule and fee constants are not yet calibrated against a live account, so a
PASSED/FAILED verdict is a diagnostic, not an authoritative answer.
Added¶
- Event-driven backtest engine with backtest/live parity against
topstep-sdk: strategies are written once against structural protocols that bothSimBrokerandAsyncTopstepClientsatisfy. SDK msgspec models (OrderModel,PositionModel,HalfTradeModel, enums,APIError) are imported verbatim, never redefined. CombineKernel— the canonical Topstep Combine rulebook as a pure state machine: two-state trailing MLL (floor ratchets only on end-of-day closed balance, breach checked every tick on realized + unrealized equity), optional daily loss limit, consistency target, position caps, 16:10 ET flatten.Backtest(bars, strategy).run()facade plusReport, and a backtesting.py-flavoured strategy dialect (SymbolStrategywithuse()indicator registration andbuy/sell/closesugar).- Tier-0 fill model over OHLCV bars: market orders rest to the next bar's open, limit orders require trade-through, one pessimistic intrabar price path shared by fills and rule-breach detection.
- Indicators via TA-Lib —
TalibIndicatordrives 152 of TA-Lib's 161 functions bar-by-bar, with typed wrappers (Sma,Ema,Rsi,Atr,Macd,BBands, …). Nothing in this package re-implements an indicator formula. - Exact-Decimal money on the instrument tick grid, FIFO lot cost basis, int-ns UTC hot path with ET session boundaries, and a 15-product instrument table.
- Data wrangling from OHLCV records or a pandas DataFrame with an explicit
stamp="open"|"close", a strict validator, and a seeded synthetic generator.
Known limitations¶
- No exchange holiday calendar. Bars on market holidays and past early-close halts are not detected or filtered anywhere. Filter them upstream.
- No continuous-contract stitching. One contract per run, inside a single front-month window. A multi-month export spanning a roll is silently merged.
- Rule and fee constants are uncalibrated — see
docs/topstep-rules.md§9. - Multi-symbol runs,
dll_enabled=True, and non-quarter-tick products (CL, GC) are not exercised end to end. - Tier-0 bar fills only. No quote, depth or MBO tiers.