harness¶
Backtest — the two-line runner that assembles the simulation stack correctly — and Report, which is what run() hands back.
harness
¶
Two-line backtest assembly: Backtest(bars, strategy).run() -> Report.
Sim-side convenience ONLY — the facade wires the exact same components a
hand-written main would (feed -> broker -> engine) and changes no semantics:
Report.result is the engine's frozen BacktestResult, byte-identical to
a hand-wired run over the same inputs. What the facade adds is assembly
correctness (AGENTS.md §§2, 6):
- ONE
TestClockshared by broker and engine. A second clock stuck at 0 would stamp every orderaccepted_ts=0(eligible for the CURRENT bar — silent look-ahead) and run session checks in 1970; the facade makes that footgun unbuildable. - Instruments derived from the feed's contract ids: the broker registers
instruments at construction only, and a bar for an unregistered contract is
a mid-run
KeyError. - Strict data validation by default: each contract's bars are checked with
ITS spec and any ERROR finding refuses to run (
validate=Falseskips; the feed's ordering invariant is enforced regardless). - Real economics always:
combine_params(account)andTopstepFees. backtesting.py's economics/semantics knobs (cash=,commission=,trade_on_close=, ...) are rejected with the project's alternative named.
Engine, broker, kernel, and strategy state are single-use, so each
Backtest runs exactly once; build a fresh Backtest (with a fresh
strategy instance) for another run.
DataValidationError
¶
DataValidationError(issues: tuple[ValidationIssue, ...])
Bases: ValueError
Bar data failed validation; issues carries every ERROR finding.
Source code in src/topstep_backtest/harness.py
Report
¶
Report(*, result: BacktestResult, stats: SummaryStats, trades: tuple[HalfTradeModel, ...], params: CombineParams, bars_gated: int | None = None, bars: tuple[Bar, ...] = (), instruments: dict[str, InstrumentSpec] | None = None, replay: Replay | None = None)
One run's full report: the untouched frozen BacktestResult, derived
SummaryStats, the broker's half-turn trade list, and the
CombineParams the run used. str(report) renders the combine
verdict, balance path, day-by-day trail, and summary stats as plain
aligned text — deterministically (pure function of the held frozen data;
no wall clock, no unordered iteration). :meth:to_html renders the same
report as a self-contained interactive tearsheet under the same
determinism contract, :meth:to_timestamped_html names that file for you,
and :meth:show opens it in the default browser.
Source code in src/topstep_backtest/harness.py
bars_gated
instance-attribute
¶
Warmup-gated bars before the strategy's first decision (None when
the strategy is not a SymbolStrategy).
bars
instance-attribute
¶
bars: tuple[Bar, ...] = bars
The OHLCV tape the run consumed (empty on hand-built reports) — what the HTML tearsheet's candlestick panes draw.
instruments
instance-attribute
¶
instruments: dict[str, InstrumentSpec] = {} if instruments is None else dict(instruments)
Specs keyed by contract id, for the tearsheet's price-axis formatting.
replay
instance-attribute
¶
replay: Replay | None = replay
The bar-by-bar recording (Backtest(..., record=True)); None
otherwise. With one present the HTML tearsheet grows the replay scrubber.
to_html
¶
to_html(path: str | PathLike[str], *, replay: ReplaySpec = 'auto', confidence: MonteCarloConfidence | None = None, crosscheck: CrossCheck | None = None) -> Path
Write the interactive HTML tearsheet to path and return it.
The document is SELF-CONTAINED — charts, styling and data are inlined, so it opens offline and can be archived next to a run. Rendering is a pure function of the held frozen data (byte-identical across calls); the named file is the only disk write.
replay controls the bar-by-bar scrubber when this report carries a
recording (Backtest(..., record=True)): "auto" embeds every
frame up to the documented limit and falls back to a loudly-labelled
window (breach-centred, else the tail) beyond it; "full" forces
every frame regardless of size; (start, end) embeds exactly that
frame range; "off" omits the scrubber. Without a recording the
knob is inert and the page says how to record one.
confidence and crosscheck add the Monte-Carlo cards (estimate
with CI, block sensitivity, year strata, bootstrap-vs-windows) — pass
the results of
:func:~topstep_backtest.metrics.confidence.mc_confidence and
:func:~topstep_backtest.metrics.confidence.crosscheck. They are
caller-computed on purpose: writing a file must never trigger
thousands of simulations as a side effect.
Source code in src/topstep_backtest/harness.py
to_timestamped_html
¶
to_timestamped_html(directory: str | PathLike[str] = '.', *, prefix: str = 'tearsheet', replay: ReplaySpec = 'auto', confidence: MonteCarloConfidence | None = None, crosscheck: CrossCheck | None = None) -> Path
Write the tearsheet to <directory>/<prefix>-<UTC stamp>.html.
The wall clock is read for the FILENAME ONLY — the document is still
the pure function of frozen run data that :meth:to_html renders, so
determinism holds where it is checked (the bytes). Missing directories
are created. A name already taken — two runs inside the same second —
gains a -2, -3, ... suffix instead of overwriting the sheet
that is already there. replay, confidence and crosscheck
are :meth:to_html's knobs, forwarded.
Source code in src/topstep_backtest/harness.py
show
¶
show(*, replay: ReplaySpec = 'auto', confidence: MonteCarloConfidence | None = None, crosscheck: CrossCheck | None = None) -> Path
Open the tearsheet in the default browser; return the file written.
This is the one API that writes without being handed a path: the
document goes to a NEW topstep-tearsheet-*.html temp file (never
overwriting anything), which is left in place so the tab survives —
delete it, or use :meth:to_html, when you want control of the path.
replay, confidence and crosscheck are :meth:to_html's
knobs, forwarded.
Source code in src/topstep_backtest/harness.py
replay_json
¶
Dump the raw :class:~topstep_backtest.replay.Replay as JSON.
The debugging tap under the scrubber's floorboards: every frame, event, and snapshot exactly as recorded, uninterpreted by any UI — for diffing two runs, or verifying what was captured before trusting a rendering of it.
Raises:
| Type | Description |
|---|---|
ValueError
|
if this report carries no recording (run with
|
Source code in src/topstep_backtest/harness.py
Backtest
¶
Backtest(data: Sequence[Bar], strategy: Strategy, *, account: AccountSize = S50K, dll_enabled: bool = False, validate: bool = True, record: bool = False, account_id: int = 1, fill_config: BarFillConfig | None = None, broker_config: SimBrokerConfig | None = None, fee_model: TopstepFees | None = None, **rejected: object)
The two-line runner: assemble the sim stack correctly and run it once.
data is a time-ordered Bar sequence (feed ordering is enforced at
construction); strategy is a bound-ready instance with typed
constructor parameters — never a class (AGENTS.md §2). Knobs
are only things that exist in this project: the account size (and its
optional Personal DLL), validation strictness, record (bar-by-bar
replay capture onto Report.replay — observation only, results are
byte-identical either way), the sim account id, and the fill/broker
fidelity configs. Economics are never knobs.
Source code in src/topstep_backtest/harness.py
from_dataframe
classmethod
¶
from_dataframe(df: Any, strategy: Strategy, *, contract_id: str, stamp: Literal['open', 'close'], unit: AggregateBarUnit, unit_number: int, account: AccountSize = S50K, dll_enabled: bool = False, validate: bool = True, record: bool = False, account_id: int = 1, fill_config: BarFillConfig | None = None, broker_config: SimBrokerConfig | None = None, fee_model: TopstepFees | None = None, **rejected: object) -> Backtest
Wrangle a pandas OHLCV DataFrame, then assemble as usual.
stamp, unit, and unit_number stay REQUIRED (no defaults),
exactly as on the wrangler: the caller must declare whether source
timestamps are bar opens or closes AND the bar span — guessing the
stamp is the classic silent one-bar look-ahead, and defaulting the
span would mis-stamp every non-1-minute bar's close (a 5-minute bar
stamped with a 60s span acts 4 minutes early against the session
clock).
Source code in src/topstep_backtest/harness.py
run
¶
run() -> Report
Run to completion synchronously (wraps asyncio.run).
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If called from inside a running event loop — use
|
Source code in src/topstep_backtest/harness.py
run_with_tearsheet
¶
run_with_tearsheet(directory: str | PathLike[str] = '.', *, prefix: str = 'tearsheet', replay: ReplaySpec = 'auto') -> tuple[Report, Path]
Run, then ALWAYS write a timestamped tearsheet: (report, path).
For "give me the HTML on every run" — exactly :meth:run followed by
:meth:Report.to_timestamped_html, so nothing about the run changes.
Inside a running event loop, compose those two yourself::
report = await backtest.arun()
written = report.to_timestamped_html("runs")
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If called from inside a running event loop (from
:meth: |
Source code in src/topstep_backtest/harness.py
arun
async
¶
arun() -> Report
Assemble fresh engine state, run once, and report.
Clock, kernel, broker, engine, and the strategy instance are all
stateful and single-use, so a Backtest refuses to run twice.