metrics.confidence¶
How much to trust the Monte-Carlo number: a double-bootstrap CI, block-length sensitivity, per-year strata, and the cross-check against real windows.
confidence
¶
How much to trust a Monte-Carlo pass probability — and when not to.
:func:~.montecarlo.monte_carlo prints one number. With 2,000 paths that
number's simulation error is negligible, which makes it dangerously easy to
read as settled — but the real uncertainty was never the path count. It is the
handful of observed trading days the bootstrap drew from, the block-length
assumption it made about how losses cluster, and the regimes the tape never
contained. This module quantifies each of those, and cross-examines the
bootstrap against the one estimator that shares none of its assumptions.
Four instruments, each aimed at a different way the point estimate lies:
- :func:
pass_probability_ci— sampling scarcity. A double bootstrap: resample the observed day set itself, re-run the Monte Carlo on each replicate, and report the 5th-95th percentile band of pass probabilities. A 65% from 500 source days and a 65% from 40 deserve different confidence, and this is the number that says so. - :func:
block_length_sensitivity— the serial-dependence assumption. Re-run the estimate across block lengths. If P(pass) swings with the block, the streak-clustering assumption is doing the work and the point estimate is fragile; if it barely moves, the choice of 5 was not load-bearing. - :func:
monte_carlo_by_year— regime blindness. Resampling cannot invent a market the tape never saw, so pooling three years into one number silently averages away the year that would have blown the account. Stratify instead: one estimate per calendar year, and the spread across them is the honest error bar non-stationarity imposes. - :func:
crosscheck— everything at once. The empirical counterpart (:func:~.windows.sequential_combines) replays real contiguous windows: no resampling, no block assumption, real regimes in real order — and pays for it in sample size. The two estimate the same quantity by opposite methods, share :func:~.montecarlo.classify_failureon purpose, and when they disagree, the disagreement IS the finding.
Every estimate here funnels through monte_carlo_from_blocks — the core the point
estimate uses — same kernel, same fill of the day blocks — so a confidence
band can never describe a different simulator than the number it qualifies.
And none of this repairs what the inputs cannot know: the rule constants are
uncalibrated (docs/topstep-rules.md §9) and a bootstrap only speaks about
futures that resemble its sample. These figures bound the statistical error;
the epistemic caveats ride along unreduced.
PassProbabilityCI
¶
Bases: Struct
A pass probability with the error bar its sample size actually earns.
point
instance-attribute
¶
The full-paths estimate on the observed days — identical to what
:func:~.montecarlo.monte_carlo reports for the same knobs and seed.
p95
instance-attribute
¶
The 5th-95th percentile band of pass probabilities across outer
resampled source-day sets: a ~90% interval for where the estimate lands
when the observed days themselves are treated as one draw from the
strategy's day distribution. Read the WIDTH before the point: a band of
30 percentage points says the tape is too short to act on, no matter how
attractive its centre. Slightly conservative by construction — each
replicate's estimate carries its own inner_paths of simulation noise,
which widens the band, never narrows it.
provisional
instance-attribute
¶
Same meaning as on :class:~.montecarlo.MonteCarloResult — and doubly
load-bearing here: under PROVISIONAL_DAY_FLOOR source days the outer
resamples are permutations of the same few observations, so even this
interval understates the true uncertainty.
BlockLengthSensitivity
¶
Bases: Struct
The same estimate at several block lengths, plus how far it moved.
results
instance-attribute
¶
results: tuple[MonteCarloResult, ...]
One full Monte Carlo per usable block length, ascending. Each row's
block_length field identifies it; length 1 is the i.i.d. resample that
destroys losing streaks and so tends to be the kindest row.
skipped_lengths
instance-attribute
¶
Requested lengths longer than the source-day count. A block longer than the tape degenerates into looping the whole tape and measures nothing, so those rows are refused rather than quietly rendered meaningless.
spread
property
¶
Max minus min pass probability across the rows.
The verdict of this instrument. Small (a few points) means the serial-dependence assumption was not load-bearing and the default block was fine. Large means the streak structure IS the result — the point estimate then deserves no more trust than the row you can least defend, and the honest report is the range, not the centre.
YearStratum
¶
Bases: Struct
One calendar year's days, bootstrapped on their own.
mc
instance-attribute
¶
mc: MonteCarloResult
source_days and provisional describe THIS year's sample; a thin
year flags itself rather than borrowing confidence from the pooled tape.
YearStratification
¶
Bases: Struct
Per-year estimates, ascending by year.
spread
property
¶
Max minus min pass probability across years; None under two
strata (a single year has no cross-regime spread to report).
This is the error bar non-stationarity imposes: the pooled estimate implicitly claims next month resembles the average of these years, and the spread says how much that claim is worth. Wider than the double-bootstrap band means regime, not sampling, is the dominant uncertainty — and no amount of paths or days fixes regime.
CrossCheckRow
¶
Bases: Struct
One outcome, estimated both ways.
outcome
instance-attribute
¶
"pass" or a :class:~.montecarlo.FailureMode name in lowercase.
null_se
instance-attribute
¶
Standard error of the window rate UNDER THE NULL that real windows
behave like bootstrap paths — binomial with the Monte-Carlo probability
over attempts windows. The null supplies the probability because the
MC side's own simulation error is negligible at thousands of paths; the
scarce side is always the windows.
divergent
instance-attribute
¶
|window_rate - mc_probability| > 2 x null_se: the real windows sit
outside what the bootstrap's own probability would produce by chance.
With ~35 windows the tolerance is naturally wide (~2 x 8 points at
p = 0.5) — a flag here is not noise, it is structure the resampling
destroyed: regime persistence past the block length, usually.
CrossCheck
¶
Bases: Struct
The bootstrap and the window sweep, forced to answer side by side.
rows
instance-attribute
¶
rows: tuple[CrossCheckRow, ...]
The four outcomes, in autopsy order: pass, MLL breach, consistency-blocked, target-not-reached.
divergent
property
¶
Any outcome divergent. Agreement is the strongest validation available without a live account: two estimators biased in opposite directions landing together. Divergence is not a bug in either — it is the finding, and the window sweep (real sequencing, real regimes) is the side to believe about WHICH days cluster.
MonteCarloConfidence
¶
Bases: Struct
The point estimate and every qualifier this module can attach to it.
ci.point == mc.pass_probability by construction (same knobs, same
seed, same deterministic core), so the bundle never shows a band around a
number it did not compute. The cross-check is absent on purpose: it needs
a window sweep, which needs the full tape and a strategy factory — run
:func:crosscheck separately when you have one.
pass_probability_ci
¶
pass_probability_ci(result: BacktestResult, *, params: CombineParams, paths: int = 2000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0, outer: int = 200, inner_paths: int = 200) -> PassProbabilityCI
Double bootstrap: a confidence band for the pass probability.
The inner bootstrap (the ordinary Monte Carlo) answers "given these
observed days, how do synthetic Combines fare?". The outer loop asks the
prior question the point estimate skips: "how different would those
observed days themselves look on another draw?" — by block-resampling the
source day set outer times (blocks, not single days: the outer draw
must preserve the same streak structure the inner one does) and re-running
the estimate on each replicate.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result
|
BacktestResult
|
The finished run whose days are the sample. |
required |
params
|
CombineParams
|
The rulebook to test against. |
required |
paths
|
int
|
Paths for the POINT estimate (matching |
2000
|
horizon_days
|
int | None
|
Sessions per attempt; default one billing month. |
None
|
block_length
|
int
|
Bootstrap block for both the outer and inner draws. |
5
|
seed
|
int
|
Drives the point estimate AND the outer replicate seeds. |
0
|
outer
|
int
|
Source-day resamples. 200 pins the 5th/95th percentiles well enough; the band's width comes from the data, not this knob. |
200
|
inner_paths
|
int
|
Paths per replicate. Deliberately smaller than |
200
|
Raises:
| Type | Description |
|---|---|
ValueError
|
on a non-positive knob or a run with no closed days. |
Source code in src/topstep_backtest/metrics/confidence.py
126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 | |
block_length_sensitivity
¶
block_length_sensitivity(result: BacktestResult, *, params: CombineParams, lengths: Sequence[int] = (1, 5, 10, 20), paths: int = 1000, horizon_days: int | None = None, seed: int = 0) -> BlockLengthSensitivity
Re-run the Monte Carlo across block lengths and report the swing.
The block length is the one genuinely arbitrary knob in the bootstrap: it
encodes how long a losing streak is assumed to travel as a unit, and no
statistic in the sample pins it. The defence is not to pick the "right"
value — there is none — but to show the estimate does not depend on it.
The same seed drives every row, so the rows differ by the assumption
under test and nothing else.
Raises:
| Type | Description |
|---|---|
ValueError
|
on empty/non-positive |
Source code in src/topstep_backtest/metrics/confidence.py
monte_carlo_by_year
¶
monte_carlo_by_year(result: BacktestResult, *, params: CombineParams, paths: int = 1000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0) -> YearStratification
One Monte Carlo per calendar year of the source run.
Stratifying by calendar year is deliberately crude: it needs no volatility model, no regime classifier to defend, and its boundaries are not chosen by looking at the outcomes. What it buys is the honest headline "P(pass) ranged from X (2024) to Y (2025)" in place of a pooled number that averages a hostile year against a kind one.
Every stratum uses the same horizon (default one billing month), so the
rows are the same question asked of different markets. On a FAILED source
run only the final year's stratum is marked source_truncated — the
breach cut that year's records short; the earlier years are complete.
Raises:
| Type | Description |
|---|---|
ValueError
|
on a run with no closed trading days. |
Source code in src/topstep_backtest/metrics/confidence.py
crosscheck
¶
crosscheck(mc: MonteCarloResult, sweep: WindowSweep) -> CrossCheck
Compare a Monte Carlo against a window sweep, outcome by outcome.
The two must share the horizon (mc.horizon_days == sweep.window_days)
— a 21-day bootstrap against 42-day windows compares answers to different
questions and would call the difference divergence. They must also share
the rulebook; that cannot be checked from here (a sweep does not carry its
params), so it is the caller's contract.
Cheap by design: both inputs are already computed, so this can run on every pair without budgeting for it.
Raises:
| Type | Description |
|---|---|
ValueError
|
on mismatched horizons or a sweep with no attempts. |
Source code in src/topstep_backtest/metrics/confidence.py
mc_confidence
¶
mc_confidence(result: BacktestResult, *, params: CombineParams, paths: int = 2000, horizon_days: int | None = None, block_length: int = 5, seed: int = 0, outer: int = 200, inner_paths: int = 200, lengths: Sequence[int] = (1, 5, 10, 20)) -> MonteCarloConfidence
One call: the estimate plus its CI, sensitivity row, and year strata.
The bundle the tearsheet renders. Sensitivity and strata run at
min(paths, 1000) paths — they are read for their spreads, which
stabilize long before the point estimate's precision is needed.
Raises:
| Type | Description |
|---|---|
ValueError
|
as the four underlying functions do. |