Skip to content

metrics.montecarlo

Block-bootstrap the run's own days through the real rule kernel for a pass probability and — more usefully — an autopsy of the failures.

montecarlo

Monte-Carlo pass probability for a Combine, by replaying resampled day sequences through the REAL rule kernel.

One backtest is one sample. It answers "did this strategy pass that tape", which is not the question worth acting on — that question is "what fraction of plausible futures does this strategy pass". This module answers the second by block-bootstrapping the observed trading days and running each synthetic sequence through a fresh CombineKernel.

Design commitments, each of which is a place a naive implementation lies:

  • The real kernel decides every verdict. Nothing here re-implements the MLL ratchet, the consistency test or the pass condition; each path drives CombineKernel through the same three calls the live engine uses. A rule fix lands here for free, and this can never drift from the engine.
  • Block bootstrap, not i.i.d. draws. Trading days are serially dependent — losing streaks cluster, and a trailing drawdown is precisely a bet against clustering. Sampling days independently would break up exactly the runs that blow accounts and would report a pass probability that is far too kind. Days are drawn in contiguous blocks (block_length) to preserve that structure.
  • Intraday excursion travels with the day. A day is not just its net P&L: the MLL is breached on realized+unrealized equity DURING the day, so each resampled day carries the worst intraday equity offset observed on the real day it came from (BacktestResult.bar_equity). Resampling closing balances alone would miss every intraday breach and badly overstate survival.
  • Deterministic. Given the same seed and inputs, identical results; random.Random is seeded per call and never the global RNG.

What it CANNOT tell you, and no amount of paths will fix:

  • The rule and fee constants underneath are uncalibrated (docs/topstep-rules.md §9). A probability computed to two decimals from unverified constants is precise, not accurate.
  • Resampling a strategy's OWN observed days assumes the future resembles the sample. It cannot invent a market regime the backtest never saw, so it understates tail risk on a short or single-regime tape.
  • Each sampled day is applied atomically as it was observed. A DLL lockout is detected and reported, but does not truncate the day's recorded P&L, because the day is a real observation rather than a re-simulation.
  • Minimum-trading-day requirements are NOT modelled — CombineParams has no such field — so "failed by not trading enough days" cannot appear in the autopsy. See docs/ROADMAP.md.

PROVISIONAL_DAY_FLOOR module-attribute

PROVISIONAL_DAY_FLOOR = 30

Observed trading days below which a bootstrap is not worth believing.

Resampling cannot create information. With a handful of source days every synthetic path is a permutation of the same few observations, so the spread of outcomes reflects the sample, not the strategy.

BILLING_MONTH_DAYS module-attribute

BILLING_MONTH_DAYS = 21

Default sessions per synthetic attempt: one subscription month.

A Combine has no time limit — it has a monthly fee, so the natural unit of "one attempt" is one fee cycle of ~21 US futures trading days, and the default question becomes the one :mod:.economics prices: is this strategy worth one more month of subscription? Pass horizon_days to model a longer commitment; lengthening it mostly converts TARGET_NOT_REACHED into whichever terminal outcome the edge actually earns.

FailureMode

Bases: IntEnum

Why a simulated path did not pass. Ordered by when it is decided.

MLL_BREACH class-attribute instance-attribute

MLL_BREACH = 1

Equity touched the trailing floor — the only TERMINAL failure the Combine rulebook has. Everything else below is a path that survived but did not qualify.

CONSISTENCY_BLOCKED class-attribute instance-attribute

CONSISTENCY_BLOCKED = 2

Cleared the profit target in dollars, but the best day was too large a share of total profit, so the kernel withheld the pass. Distinct from running out of runway: the money was made and the rule refused it.

TARGET_NOT_REACHED class-attribute instance-attribute

TARGET_NOT_REACHED = 3

Survived the horizon without reaching the profit target. The edge is too slow for the window, not too risky for it.

DayBlock

Bases: Struct

One observed trading day, as the resampling unit.

low_offset is the worst intraday equity excursion relative to the day's STARTING balance (so it is normally <= 0) — what makes an intraday MLL breach reproducible when this day is replayed at a different balance.

pnl instance-attribute

pnl: Decimal

low_offset instance-attribute

low_offset: Decimal

had_trade instance-attribute

had_trade: bool

MonteCarloResult

Bases: Struct

Outcome distribution over paths synthetic Combine attempts.

paths instance-attribute

paths: int

horizon_days instance-attribute

horizon_days: int

block_length instance-attribute

block_length: int

seed instance-attribute

seed: int

source_days instance-attribute

source_days: int

Observed trading days the bootstrap drew from. Compare against PROVISIONAL_DAY_FLOOR.

provisional instance-attribute

provisional: bool

source_days < PROVISIONAL_DAY_FLOOR: the resample has too little to work with and the spread below describes the sample, not the strategy.

source_truncated instance-attribute

source_truncated: bool

The source run ended FAILED, so its day records stop at the breach.

The sample is then survivorship-biased by construction: the days that would have followed the blow-up do not exist, so every statistic below is conditioned on the strategy having survived that long. Nothing can repair this — it is a property of the input, not of the resampling — but it must not be read past. The horizon is safe either way: it defaults to the fixed BILLING_MONTH_DAYS, never to the observed day count, which for a blown run is a survival time rather than a Combine length.

pass_probability instance-attribute

pass_probability: Decimal

mll_breach_probability instance-attribute

mll_breach_probability: Decimal

consistency_blocked_probability instance-attribute

consistency_blocked_probability: Decimal

target_not_reached_probability instance-attribute

target_not_reached_probability: Decimal

The autopsy. These four are mutually exclusive and sum to 1. They imply DIFFERENT fixes: an MLL-breach mode means resize or tighten risk; a consistency mode means the P&L is too lumpy and needs day-level throttling; a target-not-reached mode means the edge is too slow for the window and no amount of risk management helps.

dll_lock_rate instance-attribute

dll_lock_rate: Decimal

Mean daily-loss lockouts per path. Zero when no DLL is configured (the Topstep default). Not a failure — a drag.

expected_days_to_pass instance-attribute

expected_days_to_pass: Decimal | None

median_days_to_pass instance-attribute

median_days_to_pass: int | None

Days to the passing session close, over PASSING paths only. None when no path passed. Conditional on passing by construction: read it as "when it works, this is how long", never as an expected duration.

p05_ending_balance instance-attribute

p05_ending_balance: Decimal

median_ending_balance instance-attribute

median_ending_balance: Decimal

p95_ending_balance instance-attribute

p95_ending_balance: Decimal

Terminal balance percentiles across all paths (nearest-rank).

day_blocks_from

day_blocks_from(result: BacktestResult) -> tuple[DayBlock, ...]

Extract the resampling units from a finished backtest.

Each closed trading day becomes one DayBlock carrying its net P&L and its worst intraday equity excursion. Without bar_equity capture the excursion is unknown and falls back to the day's net P&L when that is a loss — a floor, not a guess: the day demonstrably reached at least its own closing loss, so the resulting breach probability is a LOWER bound.

Source code in src/topstep_backtest/metrics/montecarlo.py
def day_blocks_from(result: BacktestResult) -> tuple[DayBlock, ...]:
    """Extract the resampling units from a finished backtest.

    Each closed trading day becomes one ``DayBlock`` carrying its net P&L and
    its worst intraday equity excursion. Without ``bar_equity`` capture the
    excursion is unknown and falls back to the day's net P&L when that is a
    loss — a floor, not a guess: the day demonstrably reached at least its own
    closing loss, so the resulting breach probability is a LOWER bound.
    """
    by_day: dict[object, Decimal] = {}
    for sample in result.bar_equity:
        day = trading_day_of(sample.ts_ns)
        low = by_day.get(day)
        if low is None or sample.low < low:
            by_day[day] = sample.low

    blocks: list[DayBlock] = []
    running = result.starting_balance
    for record in result.day_records:
        observed_low = by_day.get(record.day)
        # Floor at the day's own close as well as at zero. The 16:10 flatten's
        # slippage and liquidation fees land at the session roll — AFTER the
        # last bar — so a day can close below every intraday sample. Taking the
        # observed low alone would understate that day's excursion.
        candidates = [_ZERO, record.day_pnl]
        if observed_low is not None:
            candidates.append(observed_low - running)
        blocks.append(
            DayBlock(pnl=record.day_pnl, low_offset=min(candidates), had_trade=record.had_trade)
        )
        running = record.eod_balance
    return tuple(blocks)

sample_day_path

sample_day_path(blocks: Sequence[DayBlock], horizon: int, block_length: int, rng: Random) -> list[DayBlock]

Moving-block bootstrap: contiguous runs preserve day-to-day clustering.

Source code in src/topstep_backtest/metrics/montecarlo.py
def sample_day_path(
    blocks: Sequence[DayBlock], horizon: int, block_length: int, rng: random.Random
) -> list[DayBlock]:
    """Moving-block bootstrap: contiguous runs preserve day-to-day clustering."""
    out: list[DayBlock] = []
    n = len(blocks)
    while len(out) < horizon:
        start = rng.randrange(n)
        for offset in range(block_length):
            if len(out) >= horizon:
                break
            out.append(blocks[(start + offset) % n])  # wrap: the tape is a loop
    return out

classify_failure

classify_failure(verdict: Verdict, ending: Decimal, params: CombineParams) -> FailureMode | None

Why a resolved run did not pass, or None if it did.

Shared with :func:~topstep_backtest.metrics.windows.sequential_combines ON PURPOSE. Those two estimate the same quantity by opposite methods — one resamples the observed days, the other replays real contiguous windows — and comparing them is only meaningful if a failure is attributed the same way in both. Never fork this.

Source code in src/topstep_backtest/metrics/montecarlo.py
def classify_failure(
    verdict: Verdict, ending: Decimal, params: CombineParams
) -> FailureMode | None:
    """Why a resolved run did not pass, or ``None`` if it did.

    Shared with :func:`~topstep_backtest.metrics.windows.sequential_combines`
    ON PURPOSE. Those two estimate the same quantity by opposite methods — one
    resamples the observed days, the other replays real contiguous windows —
    and comparing them is only meaningful if a failure is attributed the same
    way in both. Never fork this.
    """
    if verdict is Verdict.PASSED:
        return None
    if verdict is Verdict.FAILED:
        return FailureMode.MLL_BREACH
    # Survived but did not qualify: distinguish "made the money, rule refused"
    # from "never made the money".
    if ending >= params.starting_balance + params.profit_target:
        return FailureMode.CONSISTENCY_BLOCKED
    return FailureMode.TARGET_NOT_REACHED

nearest_rank

nearest_rank(ordered: Sequence[Decimal], pct: int) -> Decimal
Source code in src/topstep_backtest/metrics/montecarlo.py
def nearest_rank(ordered: Sequence[Decimal], pct: int) -> Decimal:
    n = len(ordered)
    rank = max(1, min(n, -(-pct * n // 100)))  # ceil, integer-only
    return ordered[rank - 1]

monte_carlo

monte_carlo(result: BacktestResult, *, params: CombineParams, paths: int = 2000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0) -> MonteCarloResult

Block-bootstrap paths synthetic Combines from a finished backtest.

Parameters:

Name Type Description Default
result BacktestResult

A completed run. Its closed trading days are the sample, and its bar_equity supplies the intraday excursions.

required
params CombineParams

The rulebook to test against — normally the same CombineParams the run used, but a different account size is a legitimate what-if.

required
paths int

Synthetic attempts to simulate.

2000
horizon_days int | None

Sessions per attempt. Defaults to BILLING_MONTH_DAYS (one subscription month) — the unit :mod:.economics bills in. Deliberately NEVER defaulted from the observed day count: a blown run's day records stop at the breach, and simulating survival-time-length Combines would flatter exactly the strategies that die fastest.

None
block_length int

Days per bootstrap block. 1 degenerates to an i.i.d. resample and will overstate the pass probability by destroying losing streaks; the default of 5 is one trading week.

5
seed int

RNG seed. Same seed and inputs give identical output.

0

Raises:

Type Description
ValueError

on a non-positive knob, or when the run closed no trading day (there is nothing to resample, and returning a confident 0% would be a lie).

Source code in src/topstep_backtest/metrics/montecarlo.py
def monte_carlo(
    result: BacktestResult,
    *,
    params: CombineParams,
    paths: int = 2000,
    horizon_days: int | None = None,
    block_length: int = 5,
    seed: int = 0,
) -> MonteCarloResult:
    """Block-bootstrap ``paths`` synthetic Combines from a finished backtest.

    Args:
        result: A completed run. Its closed trading days are the sample, and
            its ``bar_equity`` supplies the intraday excursions.
        params: The rulebook to test against — normally the same
            ``CombineParams`` the run used, but a different account size is a
            legitimate what-if.
        paths: Synthetic attempts to simulate.
        horizon_days: Sessions per attempt. Defaults to ``BILLING_MONTH_DAYS``
            (one subscription month) — the unit :mod:`.economics` bills in.
            Deliberately NEVER defaulted from the observed day count: a blown
            run's day records stop at the breach, and simulating
            survival-time-length Combines would flatter exactly the strategies
            that die fastest.
        block_length: Days per bootstrap block. 1 degenerates to an i.i.d.
            resample and will overstate the pass probability by destroying
            losing streaks; the default of 5 is one trading week.
        seed: RNG seed. Same seed and inputs give identical output.

    Raises:
        ValueError: on a non-positive knob, or when the run closed no trading
            day (there is nothing to resample, and returning a confident 0%
            would be a lie).
    """
    blocks = day_blocks_from(result)
    if not blocks:
        raise ValueError(
            "no closed trading days to resample: a Monte-Carlo over an empty "
            "sample would report a confident probability with no evidence"
        )
    return monte_carlo_from_blocks(
        blocks,
        params=params,
        paths=paths,
        horizon_days=BILLING_MONTH_DAYS if horizon_days is None else horizon_days,
        block_length=block_length,
        seed=seed,
        source_truncated=result.verdict is Verdict.FAILED,
    )

monte_carlo_from_blocks

monte_carlo_from_blocks(blocks: Sequence[DayBlock], *, params: CombineParams, paths: int, horizon_days: int, block_length: int, seed: int, source_truncated: bool) -> MonteCarloResult

The bootstrap core, over an explicit day sample.

Shared with :mod:.confidence, whose resampled-source replicates and per-year strata are day samples that never came from a single BacktestResult — every estimate there must run through THIS code or the confidence figures would describe a different simulator than the point estimate they qualify.

Source code in src/topstep_backtest/metrics/montecarlo.py
def monte_carlo_from_blocks(
    blocks: Sequence[DayBlock],
    *,
    params: CombineParams,
    paths: int,
    horizon_days: int,
    block_length: int,
    seed: int,
    source_truncated: bool,
) -> MonteCarloResult:
    """The bootstrap core, over an explicit day sample.

    Shared with :mod:`.confidence`, whose resampled-source replicates and
    per-year strata are day samples that never came from a single
    ``BacktestResult`` — every estimate there must run through THIS code or
    the confidence figures would describe a different simulator than the
    point estimate they qualify.
    """
    if paths <= 0:
        raise ValueError(f"paths must be positive, got {paths}")
    if block_length <= 0:
        raise ValueError(f"block_length must be positive, got {block_length}")
    horizon = horizon_days
    if horizon <= 0:
        raise ValueError(f"horizon_days must be positive, got {horizon}")

    rng = random.Random(seed)
    passed = 0
    modes: dict[FailureMode, int] = dict.fromkeys(FailureMode, 0)
    days_to_pass: list[int] = []
    endings: list[Decimal] = []
    total_locks = 0

    for _ in range(paths):
        days = sample_day_path(blocks, horizon, block_length, rng)
        verdict, used, locks = _run_path(days, params)
        total_locks += locks
        ending = params.starting_balance + sum((d.pnl for d in days[:used]), _ZERO)
        endings.append(ending)
        mode = classify_failure(verdict, ending, params)
        if mode is None:
            passed += 1
            days_to_pass.append(used)
        else:
            modes[mode] += 1

    endings.sort()
    n = Decimal(paths)
    return MonteCarloResult(
        paths=paths,
        horizon_days=horizon,
        block_length=block_length,
        seed=seed,
        source_days=len(blocks),
        provisional=len(blocks) < PROVISIONAL_DAY_FLOOR,
        source_truncated=source_truncated,
        pass_probability=Decimal(passed) / n,
        mll_breach_probability=Decimal(modes[FailureMode.MLL_BREACH]) / n,
        consistency_blocked_probability=Decimal(modes[FailureMode.CONSISTENCY_BLOCKED]) / n,
        target_not_reached_probability=Decimal(modes[FailureMode.TARGET_NOT_REACHED]) / n,
        dll_lock_rate=Decimal(total_locks) / n,
        expected_days_to_pass=(
            Decimal(sum(days_to_pass)) / len(days_to_pass) if days_to_pass else None
        ),
        median_days_to_pass=(
            sorted(days_to_pass)[len(days_to_pass) // 2] if days_to_pass else None
        ),
        p05_ending_balance=nearest_rank(endings, 5),
        median_ending_balance=nearest_rank(endings, 50),
        p95_ending_balance=nearest_rank(endings, 95),
    )