metrics.walkforward¶
optimize (which keeps every trial, not just the winner) and anchored walk_forward with an efficiency ratio.
walkforward
¶
Parameter search with the overfitting guards attached, and walk-forward.
optimize() was deliberately absent from this project for a long time, on the
grounds that maximizing over a pass/fail Combine metric is an overfitting
machine. That reasoning has not changed — what changed is that the guards now
exist. So the search here is built to hand you the evidence that it lied to you:
- :func:
optimizereturns EVERY configuration's result, not just the winner, and the per-period matrix needed by :func:~topstep_backtest.metrics.overfitting.probability_of_backtest_overfitting. A bare "best params" answer is the thing worth refusing. - :func:
walk_forwardnever scores a configuration on the data that selected it. Its headline is efficiency — how much of the in-sample edge survived out-of-sample — because that ratio, not the OOS number alone, is what tells you whether the fit was to signal or to the sample.
The two guards answer different questions and you want both. PBO asks whether your SELECTION generalises (symmetric in time, so it says nothing about decay). Walk-forward asks whether the edge SURVIVES FORWARD (directional, so it conflates a decaying edge with a bad selection rule). Neither subsumes the other.
Cost warning. A grid of G configurations over F folds runs G x (F + 1) complete backtests. This is the most expensive thing in the package by a wide margin, and it is serial by design — a parameter search that silently sampled would be worse than a slow one.
StrategyFactory
¶
Bases: Protocol
Builds a FRESH strategy per run.
Not optional and not a convenience: engine, broker, kernel and strategy state are all single-use here, so reusing an instance across configurations silently carries state between runs.
TrialResult
¶
Bases: Struct
One configuration's run. Kept for every configuration, winner or not.
daily_pnl
instance-attribute
¶
Per-day P&L — the row this trial contributes to the PBO matrix.
OptimizationResult
¶
Bases: Struct
Every configuration tried, plus the winner — in that order of emphasis.
best
property
¶
best: TrialResult
The in-sample winner. This is not evidence. Selecting a maximum
from a search guarantees a flattering number; feed
:attr:pbo_matrix to probability_of_backtest_overfitting and the
winner's daily_pnl to deflated_sharpe before believing it.
pbo_matrix
property
¶
Trials x periods, ready for CSCV. Rows are truncated to the shortest trial so the matrix is rectangular — configurations that died early (forced liquidation ends a run) otherwise produce ragged rows.
FoldResult
¶
Bases: Struct
One walk-forward step: chosen in-sample, scored out-of-sample.
efficiency
property
¶
oos / is — the share of the in-sample edge that survived.
1.0 means it held up entirely; below 1.0 is the normal, expected decay;
negative means the configuration lost money out-of-sample after being
chosen for making it in-sample, which is the signature of a fit to
noise. None when the in-sample objective is not positive: a ratio
against a loss is not interpretable and reporting one would invert the
sign of good and bad.
WalkForwardResult
¶
Bases: Struct
Anchored walk-forward: each fold re-selects on all data before it.
efficiency
property
¶
Aggregate sum(oos) / sum(is) across folds — the headline.
Summing before dividing is deliberate: averaging per-fold ratios lets a
single tiny-denominator fold dominate. None when the in-sample
total is not positive.
consistent_folds
property
¶
Folds that were profitable out-of-sample. A good aggregate built from one huge winner and three losers is not a robust strategy, and this is the cheapest way to see that.
net_pnl_objective
¶
net_pnl_objective(report: Report) -> Decimal
Default objective: net P&L after all fees.
Deliberately NOT the pass/fail verdict. Optimizing a binary outcome throws away almost all the information in a run and rewards configurations that scraped over the line on one tape — the exact failure this module exists to expose.
Source code in src/topstep_backtest/metrics/walkforward.py
split_by_day
¶
Cut the tape into parts chunks on TRADING-DAY boundaries.
Splitting mid-session would hand a fold a partial day — half a day's P&L scored as a day, and a strategy holding a position across the seam.
Source code in src/topstep_backtest/metrics/walkforward.py
optimize
¶
optimize(bars: Sequence[Bar], factory: StrategyFactory, grid: Sequence[Mapping[str, object]], *, account: AccountSize = S50K, objective: Callable[[Report], Decimal] = net_pnl_objective) -> OptimizationResult
Run every configuration in grid over bars and keep them all.
Raises:
| Type | Description |
|---|---|
ValueError
|
on an empty grid. |
Source code in src/topstep_backtest/metrics/walkforward.py
walk_forward
¶
walk_forward(bars: Sequence[Bar], factory: StrategyFactory, grid: Sequence[Mapping[str, object]], *, folds: int = 4, account: AccountSize = S50K, objective: Callable[[Report], Decimal] = net_pnl_objective) -> WalkForwardResult
Anchored walk-forward: select on everything before a fold, score on it.
The tape is cut into folds + 1 day-aligned chunks. Fold i selects the
best configuration over chunks 0..i and scores it on chunk i+1,
which it has never seen. Anchored rather than rolling because that is what
a trader re-optimizing periodically actually does — all history to date,
not a sliding window.
Runs len(grid) x folds in-sample backtests plus folds
out-of-sample ones. See the module cost warning.
Raises:
| Type | Description |
|---|---|
ValueError
|
on an empty grid, |