Skip to content

metrics.windows

Replay a long tape as consecutive independent Combine attempts — the empirical counterpart to the Monte-Carlo, biased the opposite way.

windows

Replay a long tape as a series of independent Combine attempts.

One backtest over three years answers a question nobody is asked: no trader runs a single continuous evaluation for three years. What they do is attempt a Combine, and if it resolves, attempt another. :func:sequential_combines models that — it cuts the tape into consecutive non-overlapping windows of window_days TRADING days and runs each as a completely fresh evaluation: new starting balance, new floor, new broker, new strategy instance.

This is the empirical counterpart to :func:~.montecarlo.monte_carlo, and the two are worth reading together precisely because they are biased in opposite directions. Monte Carlo resamples the observed days, so it gives many paths but destroys real sequencing beyond the block length — it cannot know that your worst three days were consecutive because they were the same news cycle. This module preserves sequencing and regime exactly, and pays for it in sample size: three years of data is about 35 independent 21-day windows, not 35,000. When they disagree, the disagreement is the finding, and the shared :func:~.montecarlo.classify_failure makes the comparison legitimate.

Two cuts of the same tape. :func:sequential_combines steps by the window length, so its attempts are disjoint and its pass rate is an honest (small) sample. :func:spaced_combines instead takes the number of periods you ask for and spreads their start days evenly across the tape, overlapping as much as the arithmetic requires — the question it answers is how sensitive is the outcome to when the attempt starts, which the disjoint cut cannot see. The price of overlap is that the periods are not independent samples: a bad week appears in several of them at once, so a rate over 40 overlapping periods carries the precision of 40 observations and the information of however many disjoint windows the tape holds. That number ships in the result (:attr:SpacedSweep.effective_independent_windows) precisely so the rate is never read with more confidence than the tape can back. A rolling one-day step remains refused: ~700 windows from three years pretends to a sample size the data does not have, with no compensating question it answers better.

On warm starts. Each window gets a fresh strategy, so its indicators would normally begin cold and gate away the window's opening bars — the strategy would be measured partly during its own transient, differently in every window. So the history immediately preceding a window is driven through SymbolStrategy.prewarm BEFORE the run, never through the engine: the engine sees only the window's own bars, and those preceding days therefore never appear in day_records as flat days inflating closed_days and dragging every daily percentile toward zero. The tape's opening windows have no such history to draw on and are reported fully_warm=False rather than quietly counted as equals.

WindowResult

Bases: Struct

One window's Combine attempt, resolved on its own merits.

index instance-attribute

index: int

first_day instance-attribute

first_day: date

last_day instance-attribute

last_day: date

Trading days, not calendar days — the boundaries the rules count in.

trading_days instance-attribute

trading_days: int

preload_bars instance-attribute

preload_bars: int

Bars fed through prewarm before the run so the indicators entered the window warm. Zero for the first window of a tape.

fully_warm instance-attribute

fully_warm: bool

Whether preload_bars reached the strategy's history_bars.

False means this window's indicators were still inside their transient at its first tradable bar, so its early signals are not the ones the same strategy would have produced mid-tape. Unavoidable at the start of any tape — there is no earlier data to warm from — and the reason this is a flag rather than a silent adjustment. Cold windows UNDER-trade, so leaving them in biases the pass rate DOWN.

Derived from what actually happened rather than from what was requested, so under warm_start=False every window reports False. A strategy that registers no indicators needs no history and is warm at zero.

verdict instance-attribute

verdict: Verdict

failure instance-attribute

failure: FailureMode | None

None when the window passed. Otherwise the same attribution the Monte Carlo uses (:func:~.montecarlo.classify_failure), so the two are comparable: MLL breach, consistency-blocked, or target-not-reached.

net_pnl instance-attribute

net_pnl: Decimal

ending_balance instance-attribute

ending_balance: Decimal

days_traded instance-attribute

days_traded: int

closed_trades instance-attribute

closed_trades: int

eod_trailing_drawdown instance-attribute

eod_trailing_drawdown: Decimal

Deepest decline from the EOD high-water mark — Topstep's MLL mechanic.

min_floor_headroom instance-attribute

min_floor_headroom: Decimal | None

Closest this attempt came to termination, or None without bar_equity capture. A window that passed with 40 dollars of headroom passed by luck.

consistency_headroom instance-attribute

consistency_headroom: Decimal

Dollar slack in the 50% rule at the end of the window. Negative means a balance that cleared the target still would not have been paid.

daily_pnl instance-attribute

daily_pnl: tuple[Decimal, ...]

Net P&L per closed trading day, in session order — the window's shape, not just its total. What a sparkline or a dispersion overlay draws, and the same per-day series walk-forward's TrialResult keeps. Can be shorter than trading_days when the attempt ended early (a breach stops the run).

passed property

passed: bool

cumulative_pnl property

cumulative_pnl: tuple[Decimal, ...]

Running sum of daily_pnl — the attempt's equity path from zero.

WindowSweep

Bases: Struct

Every window's attempt, plus the rates over them.

windows instance-attribute

windows: tuple[WindowResult, ...]

window_days instance-attribute

window_days: int

source_days instance-attribute

source_days: int

Trading days in the tape, before windowing.

trailing_days_dropped instance-attribute

trailing_days_dropped: int

Days at the END of the tape left over by the last whole window. Reported rather than absorbed: a partial window is not a Combine attempt, and silently scoring one would count a short evaluation as a failure to pass.

cold_start_windows instance-attribute

cold_start_windows: int

Windows whose indicators were not yet warm — see :attr:WindowResult.fully_warm. Non-zero is normal at the start of a tape.

attempts property

attempts: int

pass_rate property

pass_rate: Decimal | None

Share of attempts that PASSED. None with no attempts.

Read this beside attempts, which is usually a small number: three years of data is about 35 windows, so a pass rate has a standard error near 8 percentage points before anything else is considered. It is an estimate off a handful of samples, not a probability.

mll_breach_rate property

mll_breach_rate: Decimal | None

Attempts that hit the trailing floor — the only terminal failure. The fix is to size down.

consistency_blocked_rate property

consistency_blocked_rate: Decimal | None

Attempts that made the money and had it refused by the 50% rule. The fix is to throttle the outsized day, NOT to trade smaller.

target_not_reached_rate property

target_not_reached_rate: Decimal | None

Attempts that survived the window without reaching the target. The edge is too slow for this horizon, and neither sizing down nor throttling helps. Expect this to dominate on a short window — it is the ordinary outcome, not a failure signal.

to_html

to_html(path: str | PathLike[str]) -> Path

Write the sweep's HTML tearsheet to path and return it.

Self-contained and a pure function of the frozen sweep data, the same contract as Report.to_html; the named file is the only disk write.

Source code in src/topstep_backtest/metrics/windows.py
def to_html(self, path: str | os.PathLike[str]) -> Path:
    """Write the sweep's HTML tearsheet to ``path`` and return it.

    Self-contained and a pure function of the frozen sweep data, the same
    contract as ``Report.to_html``; the named file is the only disk write.
    """
    return _write_sweep_html(self, path)

show

show() -> Path

Open the sweep tearsheet in the default browser; return the file.

Writes a NEW topstep-sweep-*.html temp file (never overwriting anything) and leaves it in place so the tab survives — the same behaviour, and the same caveat, as Report.show.

Source code in src/topstep_backtest/metrics/windows.py
def show(self) -> Path:
    """Open the sweep tearsheet in the default browser; return the file.

    Writes a NEW ``topstep-sweep-*.html`` temp file (never overwriting
    anything) and leaves it in place so the tab survives — the same
    behaviour, and the same caveat, as ``Report.show``.
    """
    return _show_sweep_html(self)

SpacedSweep

Bases: Struct

A requested number of periods, start days spread evenly, overlap allowed.

The rates below are over OVERLAPPING attempts and are therefore not built from independent samples — read them beside :attr:effective_independent_windows, which is the evidence they actually rest on. The per-window results themselves need no such discount: each is a real backtest of a real contiguous period, and the spread across start days is exactly the start-date sensitivity this sweep exists to measure.

windows instance-attribute

windows: tuple[WindowResult, ...]

One attempt per requested period, in start-day order.

window_days instance-attribute

window_days: int

source_days instance-attribute

source_days: int

Trading days in the tape, before windowing.

stride_days instance-attribute

stride_days: Decimal | None

Nominal spacing between consecutive start days — (source_days - window_days) / (periods - 1). Actual starts are that value rounded to whole trading days. None for a single period, where spacing is meaningless.

periods property

periods: int

overlap_fraction property

overlap_fraction: Decimal | None

Share of a window that its neighbour also saw, from the nominal stride. High overlap is not a flaw — it is the cost of asking for more periods than the tape has disjoint windows — but it is the reason the rates below are not worth periods observations. None for a single period.

effective_independent_windows property

effective_independent_windows: int

Disjoint windows the tape could hold — source_days // window_days.

The honest sample size behind every rate on this sweep. Ten overlapping periods cut from a tape that holds two disjoint windows give a rate with the precision of ten observations and the information of two; sampling more periods narrows the reported spread without adding evidence.

cold_start_windows property

cold_start_windows: int

Periods whose indicators were not yet warm — see :attr:WindowResult.fully_warm. Non-zero is normal near the start of a tape.

pass_rate property

pass_rate: Decimal | None

Share of periods that PASSED, over overlapping attempts. None with no periods. See the class note before trusting the precision.

mll_breach_rate property

mll_breach_rate: Decimal | None

Periods that hit the trailing floor. The fix is to size down.

consistency_blocked_rate property

consistency_blocked_rate: Decimal | None

Periods that made the money and had it refused by the 50% rule. The fix is to throttle the outsized day, NOT to trade smaller.

target_not_reached_rate property

target_not_reached_rate: Decimal | None

Periods that survived without reaching the target. Expect this to dominate on a short window — it is the ordinary outcome there.

to_html

to_html(path: str | PathLike[str]) -> Path

Write the sweep's HTML tearsheet to path and return it.

Self-contained and a pure function of the frozen sweep data, the same contract as Report.to_html; the named file is the only disk write.

Source code in src/topstep_backtest/metrics/windows.py
def to_html(self, path: str | os.PathLike[str]) -> Path:
    """Write the sweep's HTML tearsheet to ``path`` and return it.

    Self-contained and a pure function of the frozen sweep data, the same
    contract as ``Report.to_html``; the named file is the only disk write.
    """
    return _write_sweep_html(self, path)

show

show() -> Path

Open the sweep tearsheet in the default browser; return the file.

Writes a NEW topstep-sweep-*.html temp file (never overwriting anything) and leaves it in place so the tab survives — the same behaviour, and the same caveat, as Report.show.

Source code in src/topstep_backtest/metrics/windows.py
def show(self) -> Path:
    """Open the sweep tearsheet in the default browser; return the file.

    Writes a NEW ``topstep-sweep-*.html`` temp file (never overwriting
    anything) and leaves it in place so the tab survives — the same
    behaviour, and the same caveat, as ``Report.show``.
    """
    return _show_sweep_html(self)

sequential_combines

sequential_combines(bars: Sequence[Bar], factory: Callable[[], Strategy], *, window_days: int, account: AccountSize = S50K, dll_enabled: bool = False, warm_start: bool = True, validate: bool = True) -> WindowSweep

Run one fresh Combine per non-overlapping window_days-day window.

factory must return a NEW strategy per call: engine, broker, kernel and strategy state are all single-use, so a reused instance would carry indicator state and position bookkeeping across windows and destroy the independence this function exists to provide.

With warm_start (the default) each window's strategy is prewarmed on the bars immediately preceding it, so its indicators are warm at the window's first tradable bar. That requires a SymbolStrategy; anything else must pass warm_start=False and accept a cold start per window.

validate runs ONCE over the whole tape rather than per window — the windows are slices of one already-checked series, and re-validating each would repeat the same scan tens of times for no new finding.

Cost is one complete backtest per window, and every bar is processed exactly once across the sweep plus the prewarm passes.

Raises:

Type Description
ValueError

if window_days < 1 or the tape holds fewer trading days than one window.

TypeError

if warm_start is set and factory does not return a SymbolStrategy.

Source code in src/topstep_backtest/metrics/windows.py
def sequential_combines(
    bars: Sequence[Bar],
    factory: Callable[[], Strategy],
    *,
    window_days: int,
    account: AccountSize = AccountSize.S50K,
    dll_enabled: bool = False,
    warm_start: bool = True,
    validate: bool = True,
) -> WindowSweep:
    """Run one fresh Combine per non-overlapping ``window_days``-day window.

    ``factory`` must return a NEW strategy per call: engine, broker, kernel and
    strategy state are all single-use, so a reused instance would carry
    indicator state and position bookkeeping across windows and destroy the
    independence this function exists to provide.

    With ``warm_start`` (the default) each window's strategy is prewarmed on the
    bars immediately preceding it, so its indicators are warm at the window's
    first tradable bar. That requires a ``SymbolStrategy``; anything else must
    pass ``warm_start=False`` and accept a cold start per window.

    ``validate`` runs ONCE over the whole tape rather than per window — the
    windows are slices of one already-checked series, and re-validating each
    would repeat the same scan tens of times for no new finding.

    Cost is one complete backtest per window, and every bar is processed exactly
    once across the sweep plus the prewarm passes.

    Raises:
        ValueError: if ``window_days < 1`` or the tape holds fewer trading days
            than one window.
        TypeError: if ``warm_start`` is set and ``factory`` does not return a
            ``SymbolStrategy``.
    """
    if window_days < 1:
        raise ValueError(f"window_days must be at least 1, got {window_days}")

    starts = _day_starts(bars)
    source_days = len(starts)
    if source_days < window_days:
        raise ValueError(
            f"need at least {window_days} trading days for one window, got {source_days}"
        )

    if validate:
        # One pass over the whole tape; the per-window runs then skip it. A
        # window is a slice of what was just checked, so a second scan of the
        # same bars cannot find anything the first did not.
        _backtest_cls()(bars, factory(), account=account, dll_enabled=dll_enabled, validate=True)

    params = combine_params(account, dll_enabled=dll_enabled)
    count = source_days // window_days
    results = [
        _run_window(
            bars,
            starts,
            index * window_days,
            index=index,
            window_days=window_days,
            factory=factory,
            params=params,
            account=account,
            dll_enabled=dll_enabled,
            warm_start=warm_start,
        )
        for index in range(count)
    ]

    return WindowSweep(
        windows=tuple(results),
        window_days=window_days,
        source_days=source_days,
        trailing_days_dropped=source_days - count * window_days,
        cold_start_windows=sum(1 for w in results if not w.fully_warm),
    )

spaced_combines

spaced_combines(bars: Sequence[Bar], factory: Callable[[], Strategy], *, window_days: int, periods: int, account: AccountSize = S50K, dll_enabled: bool = False, warm_start: bool = True, validate: bool = True) -> SpacedSweep

Run periods fresh Combines with start days spread evenly over the tape.

Where :func:sequential_combines lets the tape dictate how many attempts exist, this sweep lets YOU pick the number and pays for it in overlap: the first period starts on the tape's first trading day, the last starts on the last day a full window still fits, and the rest are spaced evenly between them (rounded to whole trading days). On 100 days of data, 10 periods of 40 days start roughly every 6-7 days and each shares most of its bars with its neighbours. That overlap is the point — the sweep measures how much the outcome depends on WHEN the attempt starts — but it means the periods are not independent samples; see :attr:SpacedSweep.effective_independent_windows before reading any rate as a probability.

Everything else matches :func:sequential_combines exactly (same slicing, prewarm, attribution — one shared implementation): factory must return a NEW strategy per call, warm_start prewarms each period on the bars immediately preceding it, and validate runs once over the whole tape.

Cost is one complete backtest per period, and overlapping periods re-process the shared bars once each: total work scales with periods x window_days, not with the tape length.

Raises:

Type Description
ValueError

if window_days < 1, periods < 1, the tape holds fewer trading days than one window, or periods exceeds the number of distinct start days the tape offers (which would run the same window twice and count it as two observations).

TypeError

if warm_start is set and factory does not return a SymbolStrategy.

Source code in src/topstep_backtest/metrics/windows.py
def spaced_combines(
    bars: Sequence[Bar],
    factory: Callable[[], Strategy],
    *,
    window_days: int,
    periods: int,
    account: AccountSize = AccountSize.S50K,
    dll_enabled: bool = False,
    warm_start: bool = True,
    validate: bool = True,
) -> SpacedSweep:
    """Run ``periods`` fresh Combines with start days spread evenly over the tape.

    Where :func:`sequential_combines` lets the tape dictate how many attempts
    exist, this sweep lets YOU pick the number and pays for it in overlap: the
    first period starts on the tape's first trading day, the last starts on the
    last day a full window still fits, and the rest are spaced evenly between
    them (rounded to whole trading days). On 100 days of data, 10 periods of 40
    days start roughly every 6-7 days and each shares most of its bars with its
    neighbours. That overlap is the point — the sweep measures how much the
    outcome depends on WHEN the attempt starts — but it means the periods are
    not independent samples; see :attr:`SpacedSweep.effective_independent_windows`
    before reading any rate as a probability.

    Everything else matches :func:`sequential_combines` exactly (same slicing,
    prewarm, attribution — one shared implementation): ``factory`` must return
    a NEW strategy per call, ``warm_start`` prewarms each period on the bars
    immediately preceding it, and ``validate`` runs once over the whole tape.

    Cost is one complete backtest per period, and overlapping periods re-process
    the shared bars once each: total work scales with
    ``periods x window_days``, not with the tape length.

    Raises:
        ValueError: if ``window_days < 1``, ``periods < 1``, the tape holds
            fewer trading days than one window, or ``periods`` exceeds the
            number of distinct start days the tape offers (which would run the
            same window twice and count it as two observations).
        TypeError: if ``warm_start`` is set and ``factory`` does not return a
            ``SymbolStrategy``.
    """
    if window_days < 1:
        raise ValueError(f"window_days must be at least 1, got {window_days}")
    if periods < 1:
        raise ValueError(f"periods must be at least 1, got {periods}")

    starts = _day_starts(bars)
    source_days = len(starts)
    if source_days < window_days:
        raise ValueError(
            f"need at least {window_days} trading days for one window, got {source_days}"
        )
    last_start = source_days - window_days
    if periods > last_start + 1:
        raise ValueError(
            f"periods={periods} exceeds the {last_start + 1} distinct start day(s) a "
            f"{window_days}-day window has on {source_days} trading days: the surplus "
            "periods would replay identical windows and count each as a new observation"
        )

    if validate:
        # One pass over the whole tape; the per-window runs then skip it — the
        # same reasoning as sequential_combines, and with overlap the repeated
        # scan would be even more redundant.
        _backtest_cls()(bars, factory(), account=account, dll_enabled=dll_enabled, validate=True)

    first_days: list[int]
    if periods == 1:
        first_days = [0]
        stride: Decimal | None = None
    else:
        # Evenly spaced, endpoints included, rounded half-up in exact integer
        # arithmetic. periods <= last_start + 1 guarantees the exact spacing is
        # >= 1 day, and rounding a sequence with unit-or-larger steps keeps it
        # strictly increasing — no duplicate windows can emerge here.
        denom = periods - 1
        first_days = []
        for i in range(periods):
            q, r = divmod(i * last_start, denom)
            first_days.append(q + (1 if 2 * r >= denom else 0))
        stride = Decimal(last_start) / denom

    params = combine_params(account, dll_enabled=dll_enabled)
    results = [
        _run_window(
            bars,
            starts,
            first,
            index=index,
            window_days=window_days,
            factory=factory,
            params=params,
            account=account,
            dll_enabled=dll_enabled,
            warm_start=warm_start,
        )
        for index, first in enumerate(first_days)
    ]

    return SpacedSweep(
        windows=tuple(results),
        window_days=window_days,
        source_days=source_days,
        stride_days=stride,
    )