Skip to content

metrics.confidence

How much to trust the Monte-Carlo number: a double-bootstrap CI, block-length sensitivity, per-year strata, and the cross-check against real windows.

confidence

How much to trust a Monte-Carlo pass probability — and when not to.

:func:~.montecarlo.monte_carlo prints one number. With 2,000 paths that number's simulation error is negligible, which makes it dangerously easy to read as settled — but the real uncertainty was never the path count. It is the handful of observed trading days the bootstrap drew from, the block-length assumption it made about how losses cluster, and the regimes the tape never contained. This module quantifies each of those, and cross-examines the bootstrap against the one estimator that shares none of its assumptions.

Four instruments, each aimed at a different way the point estimate lies:

  • :func:pass_probability_ci — sampling scarcity. A double bootstrap: resample the observed day set itself, re-run the Monte Carlo on each replicate, and report the 5th-95th percentile band of pass probabilities. A 65% from 500 source days and a 65% from 40 deserve different confidence, and this is the number that says so.
  • :func:block_length_sensitivity — the serial-dependence assumption. Re-run the estimate across block lengths. If P(pass) swings with the block, the streak-clustering assumption is doing the work and the point estimate is fragile; if it barely moves, the choice of 5 was not load-bearing.
  • :func:monte_carlo_by_year — regime blindness. Resampling cannot invent a market the tape never saw, so pooling three years into one number silently averages away the year that would have blown the account. Stratify instead: one estimate per calendar year, and the spread across them is the honest error bar non-stationarity imposes.
  • :func:crosscheck — everything at once. The empirical counterpart (:func:~.windows.sequential_combines) replays real contiguous windows: no resampling, no block assumption, real regimes in real order — and pays for it in sample size. The two estimate the same quantity by opposite methods, share :func:~.montecarlo.classify_failure on purpose, and when they disagree, the disagreement IS the finding.

Every estimate here funnels through monte_carlo_from_blocks — the core the point estimate uses — same kernel, same fill of the day blocks — so a confidence band can never describe a different simulator than the number it qualifies. And none of this repairs what the inputs cannot know: the rule constants are uncalibrated (docs/topstep-rules.md §9) and a bootstrap only speaks about futures that resemble its sample. These figures bound the statistical error; the epistemic caveats ride along unreduced.

PassProbabilityCI

Bases: Struct

A pass probability with the error bar its sample size actually earns.

point instance-attribute

point: Decimal

The full-paths estimate on the observed days — identical to what :func:~.montecarlo.monte_carlo reports for the same knobs and seed.

p05 instance-attribute

p05: Decimal

p95 instance-attribute

p95: Decimal

The 5th-95th percentile band of pass probabilities across outer resampled source-day sets: a ~90% interval for where the estimate lands when the observed days themselves are treated as one draw from the strategy's day distribution. Read the WIDTH before the point: a band of 30 percentage points says the tape is too short to act on, no matter how attractive its centre. Slightly conservative by construction — each replicate's estimate carries its own inner_paths of simulation noise, which widens the band, never narrows it.

outer instance-attribute

outer: int

inner_paths instance-attribute

inner_paths: int

source_days instance-attribute

source_days: int

provisional instance-attribute

provisional: bool

Same meaning as on :class:~.montecarlo.MonteCarloResult — and doubly load-bearing here: under PROVISIONAL_DAY_FLOOR source days the outer resamples are permutations of the same few observations, so even this interval understates the true uncertainty.

BlockLengthSensitivity

Bases: Struct

The same estimate at several block lengths, plus how far it moved.

results instance-attribute

results: tuple[MonteCarloResult, ...]

One full Monte Carlo per usable block length, ascending. Each row's block_length field identifies it; length 1 is the i.i.d. resample that destroys losing streaks and so tends to be the kindest row.

skipped_lengths instance-attribute

skipped_lengths: tuple[int, ...]

Requested lengths longer than the source-day count. A block longer than the tape degenerates into looping the whole tape and measures nothing, so those rows are refused rather than quietly rendered meaningless.

spread property

spread: Decimal

Max minus min pass probability across the rows.

The verdict of this instrument. Small (a few points) means the serial-dependence assumption was not load-bearing and the default block was fine. Large means the streak structure IS the result — the point estimate then deserves no more trust than the row you can least defend, and the honest report is the range, not the centre.

YearStratum

Bases: Struct

One calendar year's days, bootstrapped on their own.

year instance-attribute

year: int

mc instance-attribute

source_days and provisional describe THIS year's sample; a thin year flags itself rather than borrowing confidence from the pooled tape.

YearStratification

Bases: Struct

Per-year estimates, ascending by year.

strata instance-attribute

strata: tuple[YearStratum, ...]

spread property

spread: Decimal | None

Max minus min pass probability across years; None under two strata (a single year has no cross-regime spread to report).

This is the error bar non-stationarity imposes: the pooled estimate implicitly claims next month resembles the average of these years, and the spread says how much that claim is worth. Wider than the double-bootstrap band means regime, not sampling, is the dominant uncertainty — and no amount of paths or days fixes regime.

CrossCheckRow

Bases: Struct

One outcome, estimated both ways.

outcome instance-attribute

outcome: str

"pass" or a :class:~.montecarlo.FailureMode name in lowercase.

mc_probability instance-attribute

mc_probability: Decimal

window_rate instance-attribute

window_rate: Decimal

null_se instance-attribute

null_se: Decimal

Standard error of the window rate UNDER THE NULL that real windows behave like bootstrap paths — binomial with the Monte-Carlo probability over attempts windows. The null supplies the probability because the MC side's own simulation error is negligible at thousands of paths; the scarce side is always the windows.

divergent instance-attribute

divergent: bool

|window_rate - mc_probability| > 2 x null_se: the real windows sit outside what the bootstrap's own probability would produce by chance. With ~35 windows the tolerance is naturally wide (~2 x 8 points at p = 0.5) — a flag here is not noise, it is structure the resampling destroyed: regime persistence past the block length, usually.

CrossCheck

Bases: Struct

The bootstrap and the window sweep, forced to answer side by side.

horizon_days instance-attribute

horizon_days: int

attempts instance-attribute

attempts: int

paths instance-attribute

paths: int

rows instance-attribute

rows: tuple[CrossCheckRow, ...]

The four outcomes, in autopsy order: pass, MLL breach, consistency-blocked, target-not-reached.

divergent property

divergent: bool

Any outcome divergent. Agreement is the strongest validation available without a live account: two estimators biased in opposite directions landing together. Divergence is not a bug in either — it is the finding, and the window sweep (real sequencing, real regimes) is the side to believe about WHICH days cluster.

MonteCarloConfidence

Bases: Struct

The point estimate and every qualifier this module can attach to it.

ci.point == mc.pass_probability by construction (same knobs, same seed, same deterministic core), so the bundle never shows a band around a number it did not compute. The cross-check is absent on purpose: it needs a window sweep, which needs the full tape and a strategy factory — run :func:crosscheck separately when you have one.

mc instance-attribute

ci instance-attribute

sensitivity instance-attribute

by_year instance-attribute

pass_probability_ci

pass_probability_ci(result: BacktestResult, *, params: CombineParams, paths: int = 2000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0, outer: int = 200, inner_paths: int = 200) -> PassProbabilityCI

Double bootstrap: a confidence band for the pass probability.

The inner bootstrap (the ordinary Monte Carlo) answers "given these observed days, how do synthetic Combines fare?". The outer loop asks the prior question the point estimate skips: "how different would those observed days themselves look on another draw?" — by block-resampling the source day set outer times (blocks, not single days: the outer draw must preserve the same streak structure the inner one does) and re-running the estimate on each replicate.

Parameters:

Name Type Description Default
result BacktestResult

The finished run whose days are the sample.

required
params CombineParams

The rulebook to test against.

required
paths int

Paths for the POINT estimate (matching monte_carlo).

2000
horizon_days int | None

Sessions per attempt; default one billing month.

None
block_length int

Bootstrap block for both the outer and inner draws.

5
seed int

Drives the point estimate AND the outer replicate seeds.

0
outer int

Source-day resamples. 200 pins the 5th/95th percentiles well enough; the band's width comes from the data, not this knob.

200
inner_paths int

Paths per replicate. Deliberately smaller than paths: replicate noise only widens the band (conservative), and the outer loop multiplies whatever is spent here.

200

Raises:

Type Description
ValueError

on a non-positive knob or a run with no closed days.

Source code in src/topstep_backtest/metrics/confidence.py
def pass_probability_ci(
    result: BacktestResult,
    *,
    params: CombineParams,
    paths: int = 2000,
    horizon_days: int | None = None,
    block_length: int = 5,
    seed: int = 0,
    outer: int = 200,
    inner_paths: int = 200,
) -> PassProbabilityCI:
    """Double bootstrap: a confidence band for the pass probability.

    The inner bootstrap (the ordinary Monte Carlo) answers "given these
    observed days, how do synthetic Combines fare?". The outer loop asks the
    prior question the point estimate skips: "how different would those
    observed days themselves look on another draw?" — by block-resampling the
    source day set ``outer`` times (blocks, not single days: the outer draw
    must preserve the same streak structure the inner one does) and re-running
    the estimate on each replicate.

    Args:
        result: The finished run whose days are the sample.
        params: The rulebook to test against.
        paths: Paths for the POINT estimate (matching ``monte_carlo``).
        horizon_days: Sessions per attempt; default one billing month.
        block_length: Bootstrap block for both the outer and inner draws.
        seed: Drives the point estimate AND the outer replicate seeds.
        outer: Source-day resamples. 200 pins the 5th/95th percentiles well
            enough; the band's width comes from the data, not this knob.
        inner_paths: Paths per replicate. Deliberately smaller than ``paths``:
            replicate noise only widens the band (conservative), and the outer
            loop multiplies whatever is spent here.

    Raises:
        ValueError: on a non-positive knob or a run with no closed days.
    """
    if outer <= 0:
        raise ValueError(f"outer must be positive, got {outer}")
    if inner_paths <= 0:
        raise ValueError(f"inner_paths must be positive, got {inner_paths}")
    blocks = _blocks_or_raise(result)
    horizon = BILLING_MONTH_DAYS if horizon_days is None else horizon_days
    point = monte_carlo(
        result,
        params=params,
        paths=paths,
        horizon_days=horizon,
        block_length=block_length,
        seed=seed,
    )

    rng = random.Random(seed)
    estimates: list[Decimal] = []
    for _ in range(outer):
        replicate = tuple(sample_day_path(blocks, len(blocks), block_length, rng))
        replica_mc = monte_carlo_from_blocks(
            replicate,
            params=params,
            paths=inner_paths,
            horizon_days=horizon,
            block_length=block_length,
            seed=rng.randrange(2**32),
            source_truncated=point.source_truncated,
        )
        estimates.append(replica_mc.pass_probability)
    estimates.sort()
    return PassProbabilityCI(
        point=point.pass_probability,
        p05=nearest_rank(estimates, 5),
        p95=nearest_rank(estimates, 95),
        outer=outer,
        inner_paths=inner_paths,
        source_days=len(blocks),
        provisional=point.provisional,
    )

block_length_sensitivity

block_length_sensitivity(result: BacktestResult, *, params: CombineParams, lengths: Sequence[int] = (1, 5, 10, 20), paths: int = 1000, horizon_days: int | None = None, seed: int = 0) -> BlockLengthSensitivity

Re-run the Monte Carlo across block lengths and report the swing.

The block length is the one genuinely arbitrary knob in the bootstrap: it encodes how long a losing streak is assumed to travel as a unit, and no statistic in the sample pins it. The defence is not to pick the "right" value — there is none — but to show the estimate does not depend on it. The same seed drives every row, so the rows differ by the assumption under test and nothing else.

Raises:

Type Description
ValueError

on empty/non-positive lengths, a run with no closed days, or when every requested length exceeds the source-day count.

Source code in src/topstep_backtest/metrics/confidence.py
def block_length_sensitivity(
    result: BacktestResult,
    *,
    params: CombineParams,
    lengths: Sequence[int] = (1, 5, 10, 20),
    paths: int = 1000,
    horizon_days: int | None = None,
    seed: int = 0,
) -> BlockLengthSensitivity:
    """Re-run the Monte Carlo across block lengths and report the swing.

    The block length is the one genuinely arbitrary knob in the bootstrap: it
    encodes how long a losing streak is assumed to travel as a unit, and no
    statistic in the sample pins it. The defence is not to pick the "right"
    value — there is none — but to show the estimate does not depend on it.
    The same ``seed`` drives every row, so the rows differ by the assumption
    under test and nothing else.

    Raises:
        ValueError: on empty/non-positive ``lengths``, a run with no closed
            days, or when every requested length exceeds the source-day count.
    """
    unique = sorted(set(lengths))
    if not unique:
        raise ValueError("lengths must name at least one block length")
    if unique[0] <= 0:
        raise ValueError(f"block lengths must be positive, got {unique[0]}")
    blocks = _blocks_or_raise(result)
    usable = [length for length in unique if length <= len(blocks)]
    skipped = tuple(length for length in unique if length > len(blocks))
    if not usable:
        raise ValueError(
            f"every requested block length {unique} exceeds the {len(blocks)} "
            "source day(s) — nothing to vary"
        )
    horizon = BILLING_MONTH_DAYS if horizon_days is None else horizon_days
    truncated = result.verdict is Verdict.FAILED
    rows = tuple(
        monte_carlo_from_blocks(
            blocks,
            params=params,
            paths=paths,
            horizon_days=horizon,
            block_length=length,
            seed=seed,
            source_truncated=truncated,
        )
        for length in usable
    )
    return BlockLengthSensitivity(results=rows, skipped_lengths=skipped)

monte_carlo_by_year

monte_carlo_by_year(result: BacktestResult, *, params: CombineParams, paths: int = 1000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0) -> YearStratification

One Monte Carlo per calendar year of the source run.

Stratifying by calendar year is deliberately crude: it needs no volatility model, no regime classifier to defend, and its boundaries are not chosen by looking at the outcomes. What it buys is the honest headline "P(pass) ranged from X (2024) to Y (2025)" in place of a pooled number that averages a hostile year against a kind one.

Every stratum uses the same horizon (default one billing month), so the rows are the same question asked of different markets. On a FAILED source run only the final year's stratum is marked source_truncated — the breach cut that year's records short; the earlier years are complete.

Raises:

Type Description
ValueError

on a run with no closed trading days.

Source code in src/topstep_backtest/metrics/confidence.py
def monte_carlo_by_year(
    result: BacktestResult,
    *,
    params: CombineParams,
    paths: int = 1000,
    horizon_days: int | None = None,
    block_length: int = 5,
    seed: int = 0,
) -> YearStratification:
    """One Monte Carlo per calendar year of the source run.

    Stratifying by calendar year is deliberately crude: it needs no volatility
    model, no regime classifier to defend, and its boundaries are not chosen
    by looking at the outcomes. What it buys is the honest headline "P(pass)
    ranged from X (2024) to Y (2025)" in place of a pooled number that
    averages a hostile year against a kind one.

    Every stratum uses the same horizon (default one billing month), so the
    rows are the same question asked of different markets. On a FAILED source
    run only the final year's stratum is marked ``source_truncated`` — the
    breach cut that year's records short; the earlier years are complete.

    Raises:
        ValueError: on a run with no closed trading days.
    """
    blocks = _blocks_or_raise(result)
    horizon = BILLING_MONTH_DAYS if horizon_days is None else horizon_days
    by_year: dict[int, list[DayBlock]] = {}
    for record, block in zip(result.day_records, blocks, strict=True):
        by_year.setdefault(record.day.year, []).append(block)
    years = sorted(by_year)
    truncated_year = years[-1] if result.verdict is Verdict.FAILED else None
    return YearStratification(
        strata=tuple(
            YearStratum(
                year=year,
                mc=monte_carlo_from_blocks(
                    tuple(by_year[year]),
                    params=params,
                    paths=paths,
                    horizon_days=horizon,
                    block_length=block_length,
                    seed=seed,
                    source_truncated=year == truncated_year,
                ),
            )
            for year in years
        )
    )

crosscheck

crosscheck(mc: MonteCarloResult, sweep: WindowSweep) -> CrossCheck

Compare a Monte Carlo against a window sweep, outcome by outcome.

The two must share the horizon (mc.horizon_days == sweep.window_days) — a 21-day bootstrap against 42-day windows compares answers to different questions and would call the difference divergence. They must also share the rulebook; that cannot be checked from here (a sweep does not carry its params), so it is the caller's contract.

Cheap by design: both inputs are already computed, so this can run on every pair without budgeting for it.

Raises:

Type Description
ValueError

on mismatched horizons or a sweep with no attempts.

Source code in src/topstep_backtest/metrics/confidence.py
def crosscheck(mc: MonteCarloResult, sweep: WindowSweep) -> CrossCheck:
    """Compare a Monte Carlo against a window sweep, outcome by outcome.

    The two must share the horizon (``mc.horizon_days == sweep.window_days``)
    — a 21-day bootstrap against 42-day windows compares answers to different
    questions and would call the difference divergence. They must also share
    the rulebook; that cannot be checked from here (a sweep does not carry its
    params), so it is the caller's contract.

    Cheap by design: both inputs are already computed, so this can run on
    every pair without budgeting for it.

    Raises:
        ValueError: on mismatched horizons or a sweep with no attempts.
    """
    if mc.horizon_days != sweep.window_days:
        raise ValueError(
            f"horizons differ (monte carlo {mc.horizon_days} vs windows "
            f"{sweep.window_days}): the two would answer different questions "
            "and every gap would read as divergence"
        )
    pass_rate = sweep.pass_rate
    mll_rate = sweep.mll_breach_rate
    consistency_rate = sweep.consistency_blocked_rate
    slow_rate = sweep.target_not_reached_rate
    if pass_rate is None or mll_rate is None or consistency_rate is None or slow_rate is None:
        raise ValueError("the window sweep holds no attempts — nothing to cross-check")

    attempts = sweep.attempts
    pairs: tuple[tuple[str, Decimal, Decimal], ...] = (
        ("pass", mc.pass_probability, pass_rate),
        ("mll_breach", mc.mll_breach_probability, mll_rate),
        ("consistency_blocked", mc.consistency_blocked_probability, consistency_rate),
        ("target_not_reached", mc.target_not_reached_probability, slow_rate),
    )
    rows: list[CrossCheckRow] = []
    for outcome, mc_probability, window_rate in pairs:
        null_se = (mc_probability * (1 - mc_probability) / attempts).sqrt()
        rows.append(
            CrossCheckRow(
                outcome=outcome,
                mc_probability=mc_probability,
                window_rate=window_rate,
                null_se=null_se,
                divergent=abs(window_rate - mc_probability) > _TWO * null_se,
            )
        )
    return CrossCheck(
        horizon_days=mc.horizon_days,
        attempts=attempts,
        paths=mc.paths,
        rows=tuple(rows),
    )

mc_confidence

mc_confidence(result: BacktestResult, *, params: CombineParams, paths: int = 2000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0, outer: int = 200, inner_paths: int = 200, lengths: Sequence[int] = (1, 5, 10, 20)) -> MonteCarloConfidence

One call: the estimate plus its CI, sensitivity row, and year strata.

The bundle the tearsheet renders. Sensitivity and strata run at min(paths, 1000) paths — they are read for their spreads, which stabilize long before the point estimate's precision is needed.

Raises:

Type Description
ValueError

as the four underlying functions do.

Source code in src/topstep_backtest/metrics/confidence.py
def mc_confidence(
    result: BacktestResult,
    *,
    params: CombineParams,
    paths: int = 2000,
    horizon_days: int | None = None,
    block_length: int = 5,
    seed: int = 0,
    outer: int = 200,
    inner_paths: int = 200,
    lengths: Sequence[int] = (1, 5, 10, 20),
) -> MonteCarloConfidence:
    """One call: the estimate plus its CI, sensitivity row, and year strata.

    The bundle the tearsheet renders. Sensitivity and strata run at
    ``min(paths, 1000)`` paths — they are read for their spreads, which
    stabilize long before the point estimate's precision is needed.

    Raises:
        ValueError: as the four underlying functions do.
    """
    survey_paths = min(paths, 1000)
    return MonteCarloConfidence(
        mc=monte_carlo(
            result,
            params=params,
            paths=paths,
            horizon_days=horizon_days,
            block_length=block_length,
            seed=seed,
        ),
        ci=pass_probability_ci(
            result,
            params=params,
            paths=paths,
            horizon_days=horizon_days,
            block_length=block_length,
            seed=seed,
            outer=outer,
            inner_paths=inner_paths,
        ),
        sensitivity=block_length_sensitivity(
            result,
            params=params,
            lengths=lengths,
            paths=survey_paths,
            horizon_days=horizon_days,
            seed=seed,
        ),
        by_year=monte_carlo_by_year(
            result,
            params=params,
            paths=survey_paths,
            horizon_days=horizon_days,
            block_length=block_length,
            seed=seed,
        ),
    )