Skip to content

metrics.overfitting

The guards that make a search defensible: deflated Sharpe against a trial ledger, and PBO by CSCV.

overfitting

Deflated Sharpe Ratio — the honesty tax on how many strategies you tried.

Iterating on a strategy inside a backtester is a search, and a search over enough configurations WILL surface something with a flattering Sharpe on noise alone. The Deflated Sharpe Ratio (Bailey & Lopez de Prado, 2014) is the correction: it asks how impressive the observed Sharpe is given the number of trials it was selected from, and returns the probability that the true Sharpe is above zero.

The number that makes it work is the trial count, and a self-reported one is worthless — nobody remembers the forty parameter tweaks they abandoned. :class:TrialLedger exists so the count is recorded rather than recalled.

This module is float, deliberately. DSR is a normal-theory statistic over skew and kurtosis; there is no exact-Decimal formulation and pretending otherwise would imply a precision it does not have. Same boundary the indicator layer draws: money stays Decimal, statistics are float. Nothing here feeds a price or a P&L.

Two conventions worth stating before you compare this to anything else:

  • Per-period, NOT annualized — consistent with SummaryStats.sortino and .calmar. Annualizing a short sample manufactures confidence.
  • Dollar-Sharpe equals return-Sharpe here. Sharpe is scale-invariant, and a Combine account's base balance is fixed, so mean(pnl)/std(pnl) is identically mean(pnl/B)/std(pnl/B). Daily P&L is used directly; no base is needed and none is assumed.

DeflatedSharpe

Bases: Struct

A Sharpe ratio and what it is worth once the search is accounted for.

sharpe instance-attribute

sharpe: float

Observed per-day Sharpe. NOT annualized.

trials instance-attribute

trials: int

Independent strategy configurations the winner was selected from. The whole point of the statistic — see :class:TrialLedger.

observations instance-attribute

observations: int

Trading days behind the estimate. DSR grows with this; a great Sharpe over 12 days is not evidence.

skew instance-attribute

skew: float

kurtosis instance-attribute

kurtosis: float

Return-distribution moments (kurtosis RAW, normal = 3.0). Negative skew and fat tails both make a given Sharpe LESS impressive, and the statistic penalises them explicitly.

expected_max_sharpe instance-attribute

expected_max_sharpe: float

The Sharpe you would expect the best of trials random strategies to show under the null of no skill. If sharpe is below this, the result is worse than luck would have produced — the search alone explains it.

deflated instance-attribute

deflated: float

Probability that the true Sharpe exceeds zero, given the trial count, the sample length and the higher moments. Read as a confidence: above 0.95 is the conventional bar, below 0.5 means the observed edge is more likely an artefact of the search than a finding.

survives property

survives: bool

deflated >= 0.95 — the conventional bar, not a law.

PBOResult

Bases: Struct

Probability of Backtest Overfitting, by combinatorially symmetric cross-validation (Bailey, Borwein, Lopez de Prado & Zhu, 2015).

DSR asks "is this Sharpe impressive for the size of the search". PBO asks a sharper question about the SELECTION ITSELF: when I pick the best configuration in-sample, does it stay good out-of-sample, or was I picking noise? A strategy family can contain a genuinely good member and still score badly here if your selection rule cannot find it.

CSCV splits the observation window into splits blocks and evaluates every way of using half as in-sample and half as out-of-sample. Symmetric by construction — each split and its complement both appear — so it has no preferred direction of time, unlike walk-forward. That is a feature and a limitation: it measures selection robustness, NOT whether the edge decays, which is exactly what walk-forward is for. Run both.

pbo instance-attribute

pbo: float

Fraction of splits where the in-sample winner landed BELOW the out-of-sample median.

0.0 means the selection generalised every time; 0.5 means it did no better than picking at random, which is the signature of fitting noise. Above ~0.5 the configuration you would have chosen is actively worse than the median of the ones you rejected.

splits instance-attribute

splits: int

trials instance-attribute

trials: int

The CSCV grid: trials configurations over splits blocks, giving C(splits, splits/2) evaluations.

evaluations instance-attribute

evaluations: int

median_logit instance-attribute

median_logit: float

Median of ln(w / (1 - w)), where w is the in-sample winner's relative rank out-of-sample. Positive means the winner tends to stay above the OOS median; PBO is exactly the share of these below zero. Reported because the magnitude carries information the bare probability loses.

overfit property

overfit: bool

pbo >= 0.5 — the selection is no better than chance.

TrialLedger

TrialLedger(path: str | Path)

A file-backed count of every strategy configuration you have tried.

DSR is only as honest as its trial count, and a remembered count is always too low — the abandoned parameter sweeps are exactly the ones that inflate the winner. This records them.

Deliberately opt-in and explicit: the framework is otherwise stateless and single-use, so nothing writes to disk unless you name a path. One ledger per research question — a ledger spanning two unrelated strategies over-deflates both.

Storage is a JSON array of {"label": str, "sharpe": float | null}, append-only and human-readable on purpose: a trial count you cannot audit is no better than one you made up.

Source code in src/topstep_backtest/metrics/overfitting.py
def __init__(self, path: str | Path) -> None:
    self.path = Path(path)

path instance-attribute

path = Path(path)

record

record(label: str, *, sharpe: float | None = None) -> int

Append one trial and return the new total. Duplicate labels are kept, not deduplicated: running the same configuration twice IS two trials, and silently collapsing them would under-count the search.

Source code in src/topstep_backtest/metrics/overfitting.py
def record(self, label: str, *, sharpe: float | None = None) -> int:
    """Append one trial and return the new total. Duplicate labels are
    kept, not deduplicated: running the same configuration twice IS two
    trials, and silently collapsing them would under-count the search."""
    entries = self._load()
    entries.append({"label": label, "sharpe": sharpe})
    self.path.parent.mkdir(parents=True, exist_ok=True)
    self.path.write_text(json.dumps(entries, indent=2) + "\n")
    return len(entries)

count

count() -> int

Trials recorded so far — pass this to :func:deflated_sharpe.

Source code in src/topstep_backtest/metrics/overfitting.py
def count(self) -> int:
    """Trials recorded so far — pass this to :func:`deflated_sharpe`."""
    return len(self._load())

sharpe_variance

sharpe_variance() -> float | None

Observed variance of the recorded trial Sharpes, for deflated_sharpe(trial_sharpe_variance=...).

None with fewer than two recorded Sharpes, in which case the asymptotic fallback is used instead.

Source code in src/topstep_backtest/metrics/overfitting.py
def sharpe_variance(self) -> float | None:
    """Observed variance of the recorded trial Sharpes, for
    ``deflated_sharpe(trial_sharpe_variance=...)``.

    ``None`` with fewer than two recorded Sharpes, in which case the
    asymptotic fallback is used instead.
    """
    values: list[float] = []
    for entry in self._load():
        recorded = entry.get("sharpe")
        if isinstance(recorded, (int, float)) and not isinstance(recorded, bool):
            values.append(float(recorded))
    if len(values) < 2:
        return None
    mean = fmean(values)
    return fmean([(v - mean) ** 2 for v in values])

sharpe_ratio

sharpe_ratio(daily_pnl: Sequence[Decimal]) -> float | None

Per-day Sharpe of a daily P&L series. NOT annualized.

Uses the population standard deviation (divisor n), matching the maximum-likelihood convention the DSR derivation assumes. None when fewer than two days, or when every day is identical (zero dispersion makes the ratio undefined, not infinite).

Source code in src/topstep_backtest/metrics/overfitting.py
def sharpe_ratio(daily_pnl: Sequence[Decimal]) -> float | None:
    """Per-day Sharpe of a daily P&L series. NOT annualized.

    Uses the population standard deviation (divisor ``n``), matching the
    maximum-likelihood convention the DSR derivation assumes. ``None`` when
    fewer than two days, or when every day is identical (zero dispersion makes
    the ratio undefined, not infinite).
    """
    if len(daily_pnl) < 2:
        return None
    values = [float(p) for p in daily_pnl]
    mean = fmean(values)
    var = fmean([(v - mean) ** 2 for v in values])
    if var <= 0.0:
        return None
    return mean / math.sqrt(var)

deflated_sharpe

deflated_sharpe(daily_pnl: Sequence[Decimal], *, trials: int, trial_sharpe_variance: float | None = None) -> DeflatedSharpe | None

Deflate a strategy's Sharpe by the number of trials it was selected from.

Parameters:

Name Type Description Default
daily_pnl Sequence[Decimal]

Per-day P&L — [r.day_pnl for r in result.day_records].

required
trials int

Independent configurations tried. Counting only the configuration you kept gives trials=1, which deflates nothing and defeats the statistic. Use a :class:TrialLedger.

required
trial_sharpe_variance float | None

Variance of the Sharpes ACROSS those trials. When omitted, falls back to the asymptotic variance of a Sharpe estimate under the null, 1/(T-1). That fallback is the standard simplification and is conservative-ish, but if you have the real spread of your trial Sharpes, pass it — it is the input the deflation is most sensitive to.

None

Returns:

Type Description
DeflatedSharpe | None

None when the Sharpe itself is undefined (fewer than two days, or

DeflatedSharpe | None

zero dispersion). Never raises on a degenerate sample.

Raises:

Type Description
ValueError

if trials < 1.

Source code in src/topstep_backtest/metrics/overfitting.py
def deflated_sharpe(
    daily_pnl: Sequence[Decimal],
    *,
    trials: int,
    trial_sharpe_variance: float | None = None,
) -> DeflatedSharpe | None:
    """Deflate a strategy's Sharpe by the number of trials it was selected from.

    Args:
        daily_pnl: Per-day P&L — ``[r.day_pnl for r in result.day_records]``.
        trials: Independent configurations tried. **Counting only the
            configuration you kept gives ``trials=1``, which deflates nothing
            and defeats the statistic.** Use a :class:`TrialLedger`.
        trial_sharpe_variance: Variance of the Sharpes ACROSS those trials. When
            omitted, falls back to the asymptotic variance of a Sharpe estimate
            under the null, ``1/(T-1)``. That fallback is the standard
            simplification and is conservative-ish, but if you have the real
            spread of your trial Sharpes, pass it — it is the input the
            deflation is most sensitive to.

    Returns:
        ``None`` when the Sharpe itself is undefined (fewer than two days, or
        zero dispersion). Never raises on a degenerate sample.

    Raises:
        ValueError: if ``trials < 1``.
    """
    if trials < 1:
        raise ValueError(f"trials must be at least 1, got {trials}")
    sharpe = sharpe_ratio(daily_pnl)
    if sharpe is None:
        return None
    values = [float(p) for p in daily_pnl]
    n = len(values)
    skew, kurt = _moments(values)

    variance = 1.0 / (n - 1) if trial_sharpe_variance is None else trial_sharpe_variance
    # Expected maximum of `trials` draws from N(0, variance) — the Sharpe a
    # pure search would produce with no skill involved at all.
    if trials == 1:
        expected_max = 0.0
    else:
        gamma = _EULER_MASCHERONI
        expected_max = math.sqrt(max(variance, 0.0)) * (
            (1.0 - gamma) * _NORMAL.inv_cdf(1.0 - 1.0 / trials)
            + gamma * _NORMAL.inv_cdf(1.0 - 1.0 / (trials * math.e))
        )

    # Denominator: the standard error of the Sharpe estimate, inflated by
    # negative skew and fat tails.
    denom_sq = 1.0 - skew * sharpe + ((kurt - 1.0) / 4.0) * sharpe**2
    if denom_sq <= 0.0:
        # Pathological higher moments; the normal approximation has broken
        # down. Report no confidence rather than a complex number.
        deflated = 0.0
    else:
        z = (sharpe - expected_max) * math.sqrt(n - 1) / math.sqrt(denom_sq)
        deflated = _NORMAL.cdf(z)

    return DeflatedSharpe(
        sharpe=sharpe,
        trials=trials,
        observations=n,
        skew=skew,
        kurtosis=kurt,
        expected_max_sharpe=expected_max,
        deflated=deflated,
    )

probability_of_backtest_overfitting

probability_of_backtest_overfitting(matrix: Sequence[Sequence[float]], *, splits: int = 16) -> PBOResult

Run CSCV over a trials x observations performance matrix.

Parameters:

Name Type Description Default
matrix Sequence[Sequence[float]]

One row per configuration tried, each row that configuration's per-period P&L (or returns — Sharpe is scale-invariant, so either works as long as every row uses the same one). Every row must be the same length: they are the same periods, evaluated under different parameters.

required
splits int

Number of blocks the period axis is cut into. Must be EVEN — each evaluation uses exactly half in-sample. Bailey et al. use 16 (12,870 evaluations); lower is faster and coarser.

16

Raises:

Type Description
ValueError

on fewer than two configurations (nothing to select between, so "did the selection generalise" is not a question), an odd or too-small splits, ragged rows, or fewer observations than blocks.

Source code in src/topstep_backtest/metrics/overfitting.py
def probability_of_backtest_overfitting(
    matrix: Sequence[Sequence[float]], *, splits: int = 16
) -> PBOResult:
    """Run CSCV over a trials x observations performance matrix.

    Args:
        matrix: One row per configuration tried, each row that configuration's
            per-period P&L (or returns — Sharpe is scale-invariant, so either
            works as long as every row uses the same one). **Every row must be
            the same length**: they are the same periods, evaluated under
            different parameters.
        splits: Number of blocks the period axis is cut into. Must be EVEN —
            each evaluation uses exactly half in-sample. Bailey et al. use 16
            (12,870 evaluations); lower is faster and coarser.

    Raises:
        ValueError: on fewer than two configurations (nothing to select
            between, so "did the selection generalise" is not a question), an
            odd or too-small ``splits``, ragged rows, or fewer observations
            than blocks.
    """
    trials = len(matrix)
    if trials < 2:
        raise ValueError(f"PBO needs at least 2 configurations to select between, got {trials}")
    if splits < 2 or splits % 2 != 0:
        raise ValueError(f"splits must be even and >= 2, got {splits}")
    lengths = {len(row) for row in matrix}
    if len(lengths) != 1:
        raise ValueError(f"every configuration must cover the same periods, got lengths {lengths}")
    observations = lengths.pop()
    if observations < splits:
        raise ValueError(
            f"need at least as many observations as splits ({splits}), got {observations}"
        )

    per_trial = [_block_stats([float(v) for v in row], splits) for row in matrix]
    half = splits // 2
    logits: list[float] = []
    below = 0
    for combo in itertools.combinations(range(splits), half):
        chosen = set(combo)
        rest = [i for i in range(splits) if i not in chosen]
        is_scores = [_sharpe_from_blocks([blocks[i] for i in combo]) for blocks in per_trial]
        oos_scores = [_sharpe_from_blocks([blocks[i] for i in rest]) for blocks in per_trial]
        best = max(range(trials), key=lambda t: is_scores[t])
        # Relative rank of the IS winner among OOS scores, in (0, 1).
        rank = sum(1 for s in oos_scores if s <= oos_scores[best])
        omega = rank / (trials + 1)
        omega = min(max(omega, 1e-12), 1 - 1e-12)
        logit = math.log(omega / (1 - omega))
        logits.append(logit)
        if logit < 0:
            below += 1

    logits.sort()
    mid = len(logits) // 2
    median = logits[mid] if len(logits) % 2 else (logits[mid - 1] + logits[mid]) / 2.0
    return PBOResult(
        pbo=below / len(logits),
        splits=splits,
        trials=trials,
        evaluations=len(logits),
        median_logit=median,
    )