metrics.windows¶
Replay a long tape as consecutive independent Combine attempts — the empirical counterpart to the Monte-Carlo, biased the opposite way.
windows
¶
Replay a long tape as a series of independent Combine attempts.
One backtest over three years answers a question nobody is asked: no trader runs
a single continuous evaluation for three years. What they do is attempt a
Combine, and if it resolves, attempt another. :func:sequential_combines models
that — it cuts the tape into consecutive non-overlapping windows of
window_days TRADING days and runs each as a completely fresh evaluation:
new starting balance, new floor, new broker, new strategy instance.
This is the empirical counterpart to :func:~.montecarlo.monte_carlo, and
the two are worth reading together precisely because they are biased in opposite
directions. Monte Carlo resamples the observed days, so it gives many paths but
destroys real sequencing beyond the block length — it cannot know that your worst
three days were consecutive because they were the same news cycle. This module
preserves sequencing and regime exactly, and pays for it in sample size: three
years of data is about 35 independent 21-day windows, not 35,000. When they
disagree, the disagreement is the finding, and the shared
:func:~.montecarlo.classify_failure makes the comparison legitimate.
Two cuts of the same tape. :func:sequential_combines steps by the window
length, so its attempts are disjoint and its pass rate is an honest (small)
sample. :func:spaced_combines instead takes the number of periods you ask for
and spreads their start days evenly across the tape, overlapping as much as the
arithmetic requires — the question it answers is how sensitive is the outcome
to when the attempt starts, which the disjoint cut cannot see. The price of
overlap is that the periods are not independent samples: a bad week appears in
several of them at once, so a rate over 40 overlapping periods carries the
precision of 40 observations and the information of however many disjoint
windows the tape holds. That number ships in the result
(:attr:SpacedSweep.effective_independent_windows) precisely so the rate is
never read with more confidence than the tape can back. A rolling one-day step
remains refused: ~700 windows from three years pretends to a sample size the
data does not have, with no compensating question it answers better.
On warm starts. Each window gets a fresh strategy, so its indicators would
normally begin cold and gate away the window's opening bars — the strategy would
be measured partly during its own transient, differently in every window. So the
history immediately preceding a window is driven through
SymbolStrategy.prewarm BEFORE the run, never through the engine: the engine
sees only the window's own bars, and those preceding days therefore never appear
in day_records as flat days inflating closed_days and dragging every daily
percentile toward zero. The tape's opening windows have no such history to draw
on and are reported fully_warm=False rather than quietly counted as equals.
WindowResult
¶
Bases: Struct
One window's Combine attempt, resolved on its own merits.
last_day
instance-attribute
¶
Trading days, not calendar days — the boundaries the rules count in.
preload_bars
instance-attribute
¶
Bars fed through prewarm before the run so the indicators entered the
window warm. Zero for the first window of a tape.
fully_warm
instance-attribute
¶
Whether preload_bars reached the strategy's history_bars.
False means this window's indicators were still inside their transient
at its first tradable bar, so its early signals are not the ones the same
strategy would have produced mid-tape. Unavoidable at the start of any tape
— there is no earlier data to warm from — and the reason this is a flag
rather than a silent adjustment. Cold windows UNDER-trade, so leaving them
in biases the pass rate DOWN.
Derived from what actually happened rather than from what was requested, so
under warm_start=False every window reports False. A strategy that
registers no indicators needs no history and is warm at zero.
failure
instance-attribute
¶
failure: FailureMode | None
None when the window passed. Otherwise the same attribution the Monte
Carlo uses (:func:~.montecarlo.classify_failure), so the two are
comparable: MLL breach, consistency-blocked, or target-not-reached.
eod_trailing_drawdown
instance-attribute
¶
Deepest decline from the EOD high-water mark — Topstep's MLL mechanic.
min_floor_headroom
instance-attribute
¶
Closest this attempt came to termination, or None without bar_equity
capture. A window that passed with 40 dollars of headroom passed by luck.
consistency_headroom
instance-attribute
¶
Dollar slack in the 50% rule at the end of the window. Negative means a balance that cleared the target still would not have been paid.
daily_pnl
instance-attribute
¶
Net P&L per closed trading day, in session order — the window's shape,
not just its total. What a sparkline or a dispersion overlay draws, and the
same per-day series walk-forward's TrialResult keeps. Can be shorter
than trading_days when the attempt ended early (a breach stops the
run).
cumulative_pnl
property
¶
Running sum of daily_pnl — the attempt's equity path from zero.
WindowSweep
¶
Bases: Struct
Every window's attempt, plus the rates over them.
trailing_days_dropped
instance-attribute
¶
Days at the END of the tape left over by the last whole window. Reported rather than absorbed: a partial window is not a Combine attempt, and silently scoring one would count a short evaluation as a failure to pass.
cold_start_windows
instance-attribute
¶
Windows whose indicators were not yet warm — see
:attr:WindowResult.fully_warm. Non-zero is normal at the start of a tape.
pass_rate
property
¶
Share of attempts that PASSED. None with no attempts.
Read this beside attempts, which is usually a small number: three
years of data is about 35 windows, so a pass rate has a standard error
near 8 percentage points before anything else is considered. It is an
estimate off a handful of samples, not a probability.
mll_breach_rate
property
¶
Attempts that hit the trailing floor — the only terminal failure. The fix is to size down.
consistency_blocked_rate
property
¶
Attempts that made the money and had it refused by the 50% rule. The fix is to throttle the outsized day, NOT to trade smaller.
target_not_reached_rate
property
¶
Attempts that survived the window without reaching the target. The edge is too slow for this horizon, and neither sizing down nor throttling helps. Expect this to dominate on a short window — it is the ordinary outcome, not a failure signal.
to_html
¶
Write the sweep's HTML tearsheet to path and return it.
Self-contained and a pure function of the frozen sweep data, the same
contract as Report.to_html; the named file is the only disk write.
Source code in src/topstep_backtest/metrics/windows.py
show
¶
Open the sweep tearsheet in the default browser; return the file.
Writes a NEW topstep-sweep-*.html temp file (never overwriting
anything) and leaves it in place so the tab survives — the same
behaviour, and the same caveat, as Report.show.
Source code in src/topstep_backtest/metrics/windows.py
SpacedSweep
¶
Bases: Struct
A requested number of periods, start days spread evenly, overlap allowed.
The rates below are over OVERLAPPING attempts and are therefore not built
from independent samples — read them beside
:attr:effective_independent_windows, which is the evidence they actually
rest on. The per-window results themselves need no such discount: each is a
real backtest of a real contiguous period, and the spread across start days
is exactly the start-date sensitivity this sweep exists to measure.
windows
instance-attribute
¶
windows: tuple[WindowResult, ...]
One attempt per requested period, in start-day order.
stride_days
instance-attribute
¶
Nominal spacing between consecutive start days —
(source_days - window_days) / (periods - 1). Actual starts are that
value rounded to whole trading days. None for a single period, where
spacing is meaningless.
overlap_fraction
property
¶
Share of a window that its neighbour also saw, from the nominal
stride. High overlap is not a flaw — it is the cost of asking for more
periods than the tape has disjoint windows — but it is the reason the
rates below are not worth periods observations. None for a
single period.
effective_independent_windows
property
¶
Disjoint windows the tape could hold — source_days // window_days.
The honest sample size behind every rate on this sweep. Ten overlapping periods cut from a tape that holds two disjoint windows give a rate with the precision of ten observations and the information of two; sampling more periods narrows the reported spread without adding evidence.
cold_start_windows
property
¶
Periods whose indicators were not yet warm — see
:attr:WindowResult.fully_warm. Non-zero is normal near the start of a
tape.
pass_rate
property
¶
Share of periods that PASSED, over overlapping attempts. None
with no periods. See the class note before trusting the precision.
mll_breach_rate
property
¶
Periods that hit the trailing floor. The fix is to size down.
consistency_blocked_rate
property
¶
Periods that made the money and had it refused by the 50% rule. The fix is to throttle the outsized day, NOT to trade smaller.
target_not_reached_rate
property
¶
Periods that survived without reaching the target. Expect this to dominate on a short window — it is the ordinary outcome there.
to_html
¶
Write the sweep's HTML tearsheet to path and return it.
Self-contained and a pure function of the frozen sweep data, the same
contract as Report.to_html; the named file is the only disk write.
Source code in src/topstep_backtest/metrics/windows.py
show
¶
Open the sweep tearsheet in the default browser; return the file.
Writes a NEW topstep-sweep-*.html temp file (never overwriting
anything) and leaves it in place so the tab survives — the same
behaviour, and the same caveat, as Report.show.
Source code in src/topstep_backtest/metrics/windows.py
sequential_combines
¶
sequential_combines(bars: Sequence[Bar], factory: Callable[[], Strategy], *, window_days: int, account: AccountSize = S50K, dll_enabled: bool = False, warm_start: bool = True, validate: bool = True) -> WindowSweep
Run one fresh Combine per non-overlapping window_days-day window.
factory must return a NEW strategy per call: engine, broker, kernel and
strategy state are all single-use, so a reused instance would carry
indicator state and position bookkeeping across windows and destroy the
independence this function exists to provide.
With warm_start (the default) each window's strategy is prewarmed on the
bars immediately preceding it, so its indicators are warm at the window's
first tradable bar. That requires a SymbolStrategy; anything else must
pass warm_start=False and accept a cold start per window.
validate runs ONCE over the whole tape rather than per window — the
windows are slices of one already-checked series, and re-validating each
would repeat the same scan tens of times for no new finding.
Cost is one complete backtest per window, and every bar is processed exactly once across the sweep plus the prewarm passes.
Raises:
| Type | Description |
|---|---|
ValueError
|
if |
TypeError
|
if |
Source code in src/topstep_backtest/metrics/windows.py
447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 | |
spaced_combines
¶
spaced_combines(bars: Sequence[Bar], factory: Callable[[], Strategy], *, window_days: int, periods: int, account: AccountSize = S50K, dll_enabled: bool = False, warm_start: bool = True, validate: bool = True) -> SpacedSweep
Run periods fresh Combines with start days spread evenly over the tape.
Where :func:sequential_combines lets the tape dictate how many attempts
exist, this sweep lets YOU pick the number and pays for it in overlap: the
first period starts on the tape's first trading day, the last starts on the
last day a full window still fits, and the rest are spaced evenly between
them (rounded to whole trading days). On 100 days of data, 10 periods of 40
days start roughly every 6-7 days and each shares most of its bars with its
neighbours. That overlap is the point — the sweep measures how much the
outcome depends on WHEN the attempt starts — but it means the periods are
not independent samples; see :attr:SpacedSweep.effective_independent_windows
before reading any rate as a probability.
Everything else matches :func:sequential_combines exactly (same slicing,
prewarm, attribution — one shared implementation): factory must return
a NEW strategy per call, warm_start prewarms each period on the bars
immediately preceding it, and validate runs once over the whole tape.
Cost is one complete backtest per period, and overlapping periods re-process
the shared bars once each: total work scales with
periods x window_days, not with the tape length.
Raises:
| Type | Description |
|---|---|
ValueError
|
if |
TypeError
|
if |
Source code in src/topstep_backtest/metrics/windows.py
525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 | |