Skip to content

metrics.walkforward

optimize (which keeps every trial, not just the winner) and anchored walk_forward with an efficiency ratio.

walkforward

Parameter search with the overfitting guards attached, and walk-forward.

optimize() was deliberately absent from this project for a long time, on the grounds that maximizing over a pass/fail Combine metric is an overfitting machine. That reasoning has not changed — what changed is that the guards now exist. So the search here is built to hand you the evidence that it lied to you:

  • :func:optimize returns EVERY configuration's result, not just the winner, and the per-period matrix needed by :func:~topstep_backtest.metrics.overfitting.probability_of_backtest_overfitting. A bare "best params" answer is the thing worth refusing.
  • :func:walk_forward never scores a configuration on the data that selected it. Its headline is efficiency — how much of the in-sample edge survived out-of-sample — because that ratio, not the OOS number alone, is what tells you whether the fit was to signal or to the sample.

The two guards answer different questions and you want both. PBO asks whether your SELECTION generalises (symmetric in time, so it says nothing about decay). Walk-forward asks whether the edge SURVIVES FORWARD (directional, so it conflates a decaying edge with a bad selection rule). Neither subsumes the other.

Cost warning. A grid of G configurations over F folds runs G x (F + 1) complete backtests. This is the most expensive thing in the package by a wide margin, and it is serial by design — a parameter search that silently sampled would be worse than a slow one.

StrategyFactory

Bases: Protocol

Builds a FRESH strategy per run.

Not optional and not a convenience: engine, broker, kernel and strategy state are all single-use here, so reusing an instance across configurations silently carries state between runs.

TrialResult

Bases: Struct

One configuration's run. Kept for every configuration, winner or not.

params instance-attribute

params: dict[str, object]

objective instance-attribute

objective: Decimal

daily_pnl instance-attribute

daily_pnl: tuple[Decimal, ...]

Per-day P&L — the row this trial contributes to the PBO matrix.

net_pnl instance-attribute

net_pnl: Decimal

verdict_name instance-attribute

verdict_name: str

OptimizationResult

Bases: Struct

Every configuration tried, plus the winner — in that order of emphasis.

trials instance-attribute

trials: tuple[TrialResult, ...]

best_index instance-attribute

best_index: int

best property

The in-sample winner. This is not evidence. Selecting a maximum from a search guarantees a flattering number; feed :attr:pbo_matrix to probability_of_backtest_overfitting and the winner's daily_pnl to deflated_sharpe before believing it.

pbo_matrix property

pbo_matrix: tuple[tuple[float, ...], ...]

Trials x periods, ready for CSCV. Rows are truncated to the shortest trial so the matrix is rectangular — configurations that died early (forced liquidation ends a run) otherwise produce ragged rows.

FoldResult

Bases: Struct

One walk-forward step: chosen in-sample, scored out-of-sample.

fold instance-attribute

fold: int

params instance-attribute

params: dict[str, object]

in_sample_objective instance-attribute

in_sample_objective: Decimal

out_of_sample_objective instance-attribute

out_of_sample_objective: Decimal

in_sample_days instance-attribute

in_sample_days: int

out_of_sample_days instance-attribute

out_of_sample_days: int

efficiency property

efficiency: Decimal | None

oos / is — the share of the in-sample edge that survived.

1.0 means it held up entirely; below 1.0 is the normal, expected decay; negative means the configuration lost money out-of-sample after being chosen for making it in-sample, which is the signature of a fit to noise. None when the in-sample objective is not positive: a ratio against a loss is not interpretable and reporting one would invert the sign of good and bad.

WalkForwardResult

Bases: Struct

Anchored walk-forward: each fold re-selects on all data before it.

folds instance-attribute

folds: tuple[FoldResult, ...]

total_in_sample instance-attribute

total_in_sample: Decimal

total_out_of_sample instance-attribute

total_out_of_sample: Decimal

efficiency property

efficiency: Decimal | None

Aggregate sum(oos) / sum(is) across folds — the headline.

Summing before dividing is deliberate: averaging per-fold ratios lets a single tiny-denominator fold dominate. None when the in-sample total is not positive.

consistent_folds property

consistent_folds: int

Folds that were profitable out-of-sample. A good aggregate built from one huge winner and three losers is not a robust strategy, and this is the cheapest way to see that.

net_pnl_objective

net_pnl_objective(report: Report) -> Decimal

Default objective: net P&L after all fees.

Deliberately NOT the pass/fail verdict. Optimizing a binary outcome throws away almost all the information in a run and rewards configurations that scraped over the line on one tape — the exact failure this module exists to expose.

Source code in src/topstep_backtest/metrics/walkforward.py
def net_pnl_objective(report: Report) -> Decimal:
    """Default objective: net P&L after all fees.

    Deliberately NOT the pass/fail verdict. Optimizing a binary outcome throws
    away almost all the information in a run and rewards configurations that
    scraped over the line on one tape — the exact failure this module exists to
    expose.
    """
    return report.stats.net_pnl

split_by_day

split_by_day(bars: Sequence[Bar], parts: int) -> list[list[Bar]]

Cut the tape into parts chunks on TRADING-DAY boundaries.

Splitting mid-session would hand a fold a partial day — half a day's P&L scored as a day, and a strategy holding a position across the seam.

Source code in src/topstep_backtest/metrics/walkforward.py
def split_by_day(bars: Sequence[Bar], parts: int) -> list[list[Bar]]:
    """Cut the tape into ``parts`` chunks on TRADING-DAY boundaries.

    Splitting mid-session would hand a fold a partial day — half a day's P&L
    scored as a day, and a strategy holding a position across the seam.
    """
    days: list[list[Bar]] = []
    current: object = None
    for bar in bars:
        day = trading_day_of(bar.ts_init)
        if day != current:
            days.append([])
            current = day
        days[-1].append(bar)
    if len(days) < parts:
        raise ValueError(
            f"need at least {parts} trading days to split into {parts} parts, got {len(days)}"
        )
    edges = [(len(days) * i) // parts for i in range(parts + 1)]
    return [[bar for day in days[edges[i] : edges[i + 1]] for bar in day] for i in range(parts)]

optimize

optimize(bars: Sequence[Bar], factory: StrategyFactory, grid: Sequence[Mapping[str, object]], *, account: AccountSize = S50K, objective: Callable[[Report], Decimal] = net_pnl_objective) -> OptimizationResult

Run every configuration in grid over bars and keep them all.

Raises:

Type Description
ValueError

on an empty grid.

Source code in src/topstep_backtest/metrics/walkforward.py
def optimize(
    bars: Sequence[Bar],
    factory: StrategyFactory,
    grid: Sequence[Mapping[str, object]],
    *,
    account: AccountSize = AccountSize.S50K,
    objective: Callable[[Report], Decimal] = net_pnl_objective,
) -> OptimizationResult:
    """Run every configuration in ``grid`` over ``bars`` and keep them all.

    Raises:
        ValueError: on an empty grid.
    """
    if not grid:
        raise ValueError("grid is empty: there is nothing to optimize over")
    trials: list[TrialResult] = []
    for params in grid:
        report = _backtest_cls()(bars, factory(params), account=account).run()
        trials.append(
            TrialResult(
                params=dict(params),
                objective=objective(report),
                daily_pnl=tuple(r.day_pnl for r in report.result.day_records),
                net_pnl=report.stats.net_pnl,
                verdict_name=report.result.verdict.name,
            )
        )
    best = max(range(len(trials)), key=lambda i: trials[i].objective)
    return OptimizationResult(trials=tuple(trials), best_index=best)

walk_forward

walk_forward(bars: Sequence[Bar], factory: StrategyFactory, grid: Sequence[Mapping[str, object]], *, folds: int = 4, account: AccountSize = S50K, objective: Callable[[Report], Decimal] = net_pnl_objective) -> WalkForwardResult

Anchored walk-forward: select on everything before a fold, score on it.

The tape is cut into folds + 1 day-aligned chunks. Fold i selects the best configuration over chunks 0..i and scores it on chunk i+1, which it has never seen. Anchored rather than rolling because that is what a trader re-optimizing periodically actually does — all history to date, not a sliding window.

Runs len(grid) x folds in-sample backtests plus folds out-of-sample ones. See the module cost warning.

Raises:

Type Description
ValueError

on an empty grid, folds < 1, or too few trading days.

Source code in src/topstep_backtest/metrics/walkforward.py
def walk_forward(
    bars: Sequence[Bar],
    factory: StrategyFactory,
    grid: Sequence[Mapping[str, object]],
    *,
    folds: int = 4,
    account: AccountSize = AccountSize.S50K,
    objective: Callable[[Report], Decimal] = net_pnl_objective,
) -> WalkForwardResult:
    """Anchored walk-forward: select on everything before a fold, score on it.

    The tape is cut into ``folds + 1`` day-aligned chunks. Fold *i* selects the
    best configuration over chunks ``0..i`` and scores it on chunk ``i+1``,
    which it has never seen. Anchored rather than rolling because that is what
    a trader re-optimizing periodically actually does — all history to date,
    not a sliding window.

    Runs ``len(grid) x folds`` in-sample backtests plus ``folds``
    out-of-sample ones. See the module cost warning.

    Raises:
        ValueError: on an empty grid, ``folds < 1``, or too few trading days.
    """
    if not grid:
        raise ValueError("grid is empty: there is nothing to optimize over")
    if folds < 1:
        raise ValueError(f"folds must be at least 1, got {folds}")

    chunks = split_by_day(bars, folds + 1)
    results: list[FoldResult] = []
    total_is = _ZERO
    total_oos = _ZERO
    for i in range(folds):
        in_sample = [bar for chunk in chunks[: i + 1] for bar in chunk]
        out_of_sample = chunks[i + 1]
        chosen = optimize(in_sample, factory, grid, account=account, objective=objective)
        winner = chosen.best
        oos_report = _backtest_cls()(out_of_sample, factory(winner.params), account=account).run()
        oos_objective = objective(oos_report)
        total_is += winner.objective
        total_oos += oos_objective
        results.append(
            FoldResult(
                fold=i,
                params=dict(winner.params),
                in_sample_objective=winner.objective,
                out_of_sample_objective=oos_objective,
                in_sample_days=len(chosen.best.daily_pnl),
                out_of_sample_days=len(oos_report.result.day_records),
            )
        )
    return WalkForwardResult(
        folds=tuple(results), total_in_sample=total_is, total_out_of_sample=total_oos
    )