metrics.montecarlo¶
Block-bootstrap the run's own days through the real rule kernel for a pass probability and — more usefully — an autopsy of the failures.
montecarlo
¶
Monte-Carlo pass probability for a Combine, by replaying resampled day sequences through the REAL rule kernel.
One backtest is one sample. It answers "did this strategy pass that tape",
which is not the question worth acting on — that question is "what fraction of
plausible futures does this strategy pass". This module answers the second by
block-bootstrapping the observed trading days and running each synthetic
sequence through a fresh CombineKernel.
Design commitments, each of which is a place a naive implementation lies:
- The real kernel decides every verdict. Nothing here re-implements the
MLL ratchet, the consistency test or the pass condition; each path drives
CombineKernelthrough the same three calls the live engine uses. A rule fix lands here for free, and this can never drift from the engine. - Block bootstrap, not i.i.d. draws. Trading days are serially dependent —
losing streaks cluster, and a trailing drawdown is precisely a bet against
clustering. Sampling days independently would break up exactly the runs that
blow accounts and would report a pass probability that is far too kind. Days
are drawn in contiguous blocks (
block_length) to preserve that structure. - Intraday excursion travels with the day. A day is not just its net P&L:
the MLL is breached on realized+unrealized equity DURING the day, so each
resampled day carries the worst intraday equity offset observed on the real
day it came from (
BacktestResult.bar_equity). Resampling closing balances alone would miss every intraday breach and badly overstate survival. - Deterministic. Given the same
seedand inputs, identical results;random.Randomis seeded per call and never the global RNG.
What it CANNOT tell you, and no amount of paths will fix:
- The rule and fee constants underneath are uncalibrated (docs/topstep-rules.md §9). A probability computed to two decimals from unverified constants is precise, not accurate.
- Resampling a strategy's OWN observed days assumes the future resembles the sample. It cannot invent a market regime the backtest never saw, so it understates tail risk on a short or single-regime tape.
- Each sampled day is applied atomically as it was observed. A DLL lockout is detected and reported, but does not truncate the day's recorded P&L, because the day is a real observation rather than a re-simulation.
- Minimum-trading-day requirements are NOT modelled —
CombineParamshas no such field — so "failed by not trading enough days" cannot appear in the autopsy. See docs/ROADMAP.md.
PROVISIONAL_DAY_FLOOR
module-attribute
¶
Observed trading days below which a bootstrap is not worth believing.
Resampling cannot create information. With a handful of source days every synthetic path is a permutation of the same few observations, so the spread of outcomes reflects the sample, not the strategy.
BILLING_MONTH_DAYS
module-attribute
¶
Default sessions per synthetic attempt: one subscription month.
A Combine has no time limit — it has a monthly fee, so the natural unit of
"one attempt" is one fee cycle of ~21 US futures trading days, and the default
question becomes the one :mod:.economics prices: is this strategy worth one
more month of subscription? Pass horizon_days to model a longer
commitment; lengthening it mostly converts TARGET_NOT_REACHED into whichever
terminal outcome the edge actually earns.
FailureMode
¶
Bases: IntEnum
Why a simulated path did not pass. Ordered by when it is decided.
MLL_BREACH
class-attribute
instance-attribute
¶
Equity touched the trailing floor — the only TERMINAL failure the Combine rulebook has. Everything else below is a path that survived but did not qualify.
CONSISTENCY_BLOCKED
class-attribute
instance-attribute
¶
Cleared the profit target in dollars, but the best day was too large a share of total profit, so the kernel withheld the pass. Distinct from running out of runway: the money was made and the rule refused it.
TARGET_NOT_REACHED
class-attribute
instance-attribute
¶
Survived the horizon without reaching the profit target. The edge is too slow for the window, not too risky for it.
DayBlock
¶
Bases: Struct
One observed trading day, as the resampling unit.
low_offset is the worst intraday equity excursion relative to the day's
STARTING balance (so it is normally <= 0) — what makes an intraday MLL
breach reproducible when this day is replayed at a different balance.
MonteCarloResult
¶
Bases: Struct
Outcome distribution over paths synthetic Combine attempts.
source_days
instance-attribute
¶
Observed trading days the bootstrap drew from. Compare against
PROVISIONAL_DAY_FLOOR.
provisional
instance-attribute
¶
source_days < PROVISIONAL_DAY_FLOOR: the resample has too little to
work with and the spread below describes the sample, not the strategy.
source_truncated
instance-attribute
¶
The source run ended FAILED, so its day records stop at the breach.
The sample is then survivorship-biased by construction: the days that
would have followed the blow-up do not exist, so every statistic below is
conditioned on the strategy having survived that long. Nothing can repair
this — it is a property of the input, not of the resampling — but it must
not be read past. The horizon is safe either way: it defaults to the fixed
BILLING_MONTH_DAYS, never to the observed day count, which for a blown
run is a survival time rather than a Combine length.
target_not_reached_probability
instance-attribute
¶
The autopsy. These four are mutually exclusive and sum to 1. They imply DIFFERENT fixes: an MLL-breach mode means resize or tighten risk; a consistency mode means the P&L is too lumpy and needs day-level throttling; a target-not-reached mode means the edge is too slow for the window and no amount of risk management helps.
dll_lock_rate
instance-attribute
¶
Mean daily-loss lockouts per path. Zero when no DLL is configured (the Topstep default). Not a failure — a drag.
median_days_to_pass
instance-attribute
¶
Days to the passing session close, over PASSING paths only. None
when no path passed. Conditional on passing by construction: read it as
"when it works, this is how long", never as an expected duration.
p95_ending_balance
instance-attribute
¶
Terminal balance percentiles across all paths (nearest-rank).
day_blocks_from
¶
day_blocks_from(result: BacktestResult) -> tuple[DayBlock, ...]
Extract the resampling units from a finished backtest.
Each closed trading day becomes one DayBlock carrying its net P&L and
its worst intraday equity excursion. Without bar_equity capture the
excursion is unknown and falls back to the day's net P&L when that is a
loss — a floor, not a guess: the day demonstrably reached at least its own
closing loss, so the resulting breach probability is a LOWER bound.
Source code in src/topstep_backtest/metrics/montecarlo.py
sample_day_path
¶
sample_day_path(blocks: Sequence[DayBlock], horizon: int, block_length: int, rng: Random) -> list[DayBlock]
Moving-block bootstrap: contiguous runs preserve day-to-day clustering.
Source code in src/topstep_backtest/metrics/montecarlo.py
classify_failure
¶
classify_failure(verdict: Verdict, ending: Decimal, params: CombineParams) -> FailureMode | None
Why a resolved run did not pass, or None if it did.
Shared with :func:~topstep_backtest.metrics.windows.sequential_combines
ON PURPOSE. Those two estimate the same quantity by opposite methods — one
resamples the observed days, the other replays real contiguous windows —
and comparing them is only meaningful if a failure is attributed the same
way in both. Never fork this.
Source code in src/topstep_backtest/metrics/montecarlo.py
nearest_rank
¶
monte_carlo
¶
monte_carlo(result: BacktestResult, *, params: CombineParams, paths: int = 2000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0) -> MonteCarloResult
Block-bootstrap paths synthetic Combines from a finished backtest.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
BacktestResult
|
A completed run. Its closed trading days are the sample, and
its |
required |
params
|
CombineParams
|
The rulebook to test against — normally the same
|
required |
paths
|
int
|
Synthetic attempts to simulate. |
2000
|
horizon_days
|
int | None
|
Sessions per attempt. Defaults to |
None
|
block_length
|
int
|
Days per bootstrap block. 1 degenerates to an i.i.d. resample and will overstate the pass probability by destroying losing streaks; the default of 5 is one trading week. |
5
|
seed
|
int
|
RNG seed. Same seed and inputs give identical output. |
0
|
Raises:
| Type | Description |
|---|---|
ValueError
|
on a non-positive knob, or when the run closed no trading day (there is nothing to resample, and returning a confident 0% would be a lie). |
Source code in src/topstep_backtest/metrics/montecarlo.py
monte_carlo_from_blocks
¶
monte_carlo_from_blocks(blocks: Sequence[DayBlock], *, params: CombineParams, paths: int, horizon_days: int, block_length: int, seed: int, source_truncated: bool) -> MonteCarloResult
The bootstrap core, over an explicit day sample.
Shared with :mod:.confidence, whose resampled-source replicates and
per-year strata are day samples that never came from a single
BacktestResult — every estimate there must run through THIS code or
the confidence figures would describe a different simulator than the
point estimate they qualify.