metrics.overfitting¶
The guards that make a search defensible: deflated Sharpe against a trial ledger, and PBO by CSCV.
overfitting
¶
Deflated Sharpe Ratio — the honesty tax on how many strategies you tried.
Iterating on a strategy inside a backtester is a search, and a search over enough configurations WILL surface something with a flattering Sharpe on noise alone. The Deflated Sharpe Ratio (Bailey & Lopez de Prado, 2014) is the correction: it asks how impressive the observed Sharpe is given the number of trials it was selected from, and returns the probability that the true Sharpe is above zero.
The number that makes it work is the trial count, and a self-reported one is
worthless — nobody remembers the forty parameter tweaks they abandoned.
:class:TrialLedger exists so the count is recorded rather than recalled.
This module is float, deliberately. DSR is a normal-theory statistic over
skew and kurtosis; there is no exact-Decimal formulation and pretending
otherwise would imply a precision it does not have. Same boundary the indicator
layer draws: money stays Decimal, statistics are float. Nothing here feeds a
price or a P&L.
Two conventions worth stating before you compare this to anything else:
- Per-period, NOT annualized — consistent with
SummaryStats.sortinoand.calmar. Annualizing a short sample manufactures confidence. - Dollar-Sharpe equals return-Sharpe here. Sharpe is scale-invariant, and a
Combine account's base balance is fixed, so
mean(pnl)/std(pnl)is identicallymean(pnl/B)/std(pnl/B). Daily P&L is used directly; no base is needed and none is assumed.
DeflatedSharpe
¶
Bases: Struct
A Sharpe ratio and what it is worth once the search is accounted for.
trials
instance-attribute
¶
Independent strategy configurations the winner was selected from. The
whole point of the statistic — see :class:TrialLedger.
observations
instance-attribute
¶
Trading days behind the estimate. DSR grows with this; a great Sharpe over 12 days is not evidence.
kurtosis
instance-attribute
¶
Return-distribution moments (kurtosis RAW, normal = 3.0). Negative skew and fat tails both make a given Sharpe LESS impressive, and the statistic penalises them explicitly.
expected_max_sharpe
instance-attribute
¶
The Sharpe you would expect the best of trials random strategies to
show under the null of no skill. If sharpe is below this, the result
is worse than luck would have produced — the search alone explains it.
deflated
instance-attribute
¶
Probability that the true Sharpe exceeds zero, given the trial count, the sample length and the higher moments. Read as a confidence: above 0.95 is the conventional bar, below 0.5 means the observed edge is more likely an artefact of the search than a finding.
PBOResult
¶
Bases: Struct
Probability of Backtest Overfitting, by combinatorially symmetric cross-validation (Bailey, Borwein, Lopez de Prado & Zhu, 2015).
DSR asks "is this Sharpe impressive for the size of the search". PBO asks a sharper question about the SELECTION ITSELF: when I pick the best configuration in-sample, does it stay good out-of-sample, or was I picking noise? A strategy family can contain a genuinely good member and still score badly here if your selection rule cannot find it.
CSCV splits the observation window into splits blocks and evaluates
every way of using half as in-sample and half as out-of-sample. Symmetric
by construction — each split and its complement both appear — so it has no
preferred direction of time, unlike walk-forward. That is a feature and a
limitation: it measures selection robustness, NOT whether the edge decays,
which is exactly what walk-forward is for. Run both.
pbo
instance-attribute
¶
Fraction of splits where the in-sample winner landed BELOW the out-of-sample median.
0.0 means the selection generalised every time; 0.5 means it did no better than picking at random, which is the signature of fitting noise. Above ~0.5 the configuration you would have chosen is actively worse than the median of the ones you rejected.
trials
instance-attribute
¶
The CSCV grid: trials configurations over splits blocks, giving
C(splits, splits/2) evaluations.
median_logit
instance-attribute
¶
Median of ln(w / (1 - w)), where w is the in-sample winner's
relative rank out-of-sample. Positive means the winner tends to stay above
the OOS median; PBO is exactly the share of these below zero. Reported
because the magnitude carries information the bare probability loses.
TrialLedger
¶
A file-backed count of every strategy configuration you have tried.
DSR is only as honest as its trial count, and a remembered count is always too low — the abandoned parameter sweeps are exactly the ones that inflate the winner. This records them.
Deliberately opt-in and explicit: the framework is otherwise stateless and single-use, so nothing writes to disk unless you name a path. One ledger per research question — a ledger spanning two unrelated strategies over-deflates both.
Storage is a JSON array of {"label": str, "sharpe": float | null},
append-only and human-readable on purpose: a trial count you cannot audit
is no better than one you made up.
Source code in src/topstep_backtest/metrics/overfitting.py
record
¶
Append one trial and return the new total. Duplicate labels are kept, not deduplicated: running the same configuration twice IS two trials, and silently collapsing them would under-count the search.
Source code in src/topstep_backtest/metrics/overfitting.py
count
¶
sharpe_variance
¶
Observed variance of the recorded trial Sharpes, for
deflated_sharpe(trial_sharpe_variance=...).
None with fewer than two recorded Sharpes, in which case the
asymptotic fallback is used instead.
Source code in src/topstep_backtest/metrics/overfitting.py
sharpe_ratio
¶
Per-day Sharpe of a daily P&L series. NOT annualized.
Uses the population standard deviation (divisor n), matching the
maximum-likelihood convention the DSR derivation assumes. None when
fewer than two days, or when every day is identical (zero dispersion makes
the ratio undefined, not infinite).
Source code in src/topstep_backtest/metrics/overfitting.py
deflated_sharpe
¶
deflated_sharpe(daily_pnl: Sequence[Decimal], *, trials: int, trial_sharpe_variance: float | None = None) -> DeflatedSharpe | None
Deflate a strategy's Sharpe by the number of trials it was selected from.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
daily_pnl
|
Sequence[Decimal]
|
Per-day P&L — |
required |
trials
|
int
|
Independent configurations tried. Counting only the
configuration you kept gives |
required |
trial_sharpe_variance
|
float | None
|
Variance of the Sharpes ACROSS those trials. When
omitted, falls back to the asymptotic variance of a Sharpe estimate
under the null, |
None
|
Returns:
| Type | Description |
|---|---|
DeflatedSharpe | None
|
|
DeflatedSharpe | None
|
zero dispersion). Never raises on a degenerate sample. |
Raises:
| Type | Description |
|---|---|
ValueError
|
if |
Source code in src/topstep_backtest/metrics/overfitting.py
probability_of_backtest_overfitting
¶
probability_of_backtest_overfitting(matrix: Sequence[Sequence[float]], *, splits: int = 16) -> PBOResult
Run CSCV over a trials x observations performance matrix.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
matrix
|
Sequence[Sequence[float]]
|
One row per configuration tried, each row that configuration's per-period P&L (or returns — Sharpe is scale-invariant, so either works as long as every row uses the same one). Every row must be the same length: they are the same periods, evaluated under different parameters. |
required |
splits
|
int
|
Number of blocks the period axis is cut into. Must be EVEN — each evaluation uses exactly half in-sample. Bailey et al. use 16 (12,870 evaluations); lower is faster and coarser. |
16
|
Raises:
| Type | Description |
|---|---|
ValueError
|
on fewer than two configurations (nothing to select
between, so "did the selection generalise" is not a question), an
odd or too-small |