metrics.stats¶
SummaryStats and its three nested blocks. Every field states its basis, because most metrics admit two honest answers.
stats
¶
Combine-centric summary statistics derived from a finished backtest run.
Every metric's basis is stated explicitly because most admit two honest bases (gross vs net of fees, bar-close vs intrabar) and mixing them silently is how reports lie:
- A "closing half-turn" is a broker trade record whose
profit_and_lossis notNoneand that is not voided.profit_and_lossis the GROSS realized P&L of the closed portion; fees and commissions are charged on EVERY half-turn (open and close) and deducted from balance separately. So per-close classification (win rate, expectancy, profit factor) is gross, while the aggregate net figure isnet_pnl(ending minus starting balance, all fees included). Re-attributing opening fees to round trips would need new FIFO pairing whose live-gateway equivalence is unverified (docs/topstep-rules.md §9) — deliberately not offered. - Drawdown is measured on the per-bar CLOSE equity curve, closed with one
terminal mark at
ending_balance: the engine's final session roll (16:10 flatten slippage + liquidation fees) lands AFTER the last curve point, and without the terminal markmax_drawdowncould sit below the net loss printed beside it. Intrabar excursions are not observable fromBacktestResultand no proxy is attempted. - The pass/fail
verdictis THREE-state (PASSED/FAILED/IN_PROGRESS) and is mirrored fromBacktestResultunchanged. Most runs endIN_PROGRESS— the tape ran out before the combine resolved — so a two-state read of it ("not passed" => failed) misreports the common case as a blowup.passedandfailedare bothFalsethere. - Consistency (docs/topstep-rules.md §4): passing requires
best_day <= consistency_pct x total_profit, wherebest_dayis the largest traded-day EOD-balance delta (net of fees, never below zero) andtotal_profitis the last closed balance minus start.consistency_pctis not aBacktestResultfield — the caller passes it from theCombineParamsthe run used.
PROVISIONAL_TRADE_FLOOR
module-attribute
¶
Closing half-turns below which every distributional metric is noise.
Not a rule of thumb this project invented and not a hard boundary — the
standard error of an expectancy estimate falls as 1/sqrt(n), so somewhere
around here a strategy's measured edge stops being distinguishable from its
sampling error. SummaryStats.provisional is True below it so a thin
run cannot be read as a verdict by accident.
DailyStats
¶
Bases: Struct
Distribution of per-trading-day P&L, for reasoning about daily limits.
Every figure is the day's EOD-balance delta (DayRecord.day_pnl): NET
of all fees, and only for CLOSED trading days — an in-progress final day
contributes nothing. Percentiles are nearest-rank (see :func:_percentile)
and are None when no day has closed.
closed_days
instance-attribute
¶
Trading days that closed — the denominator for everything below.
flat_days
instance-attribute
¶
Day counts by sign of day_pnl. A scratch day (exactly 0) is flat,
counted in neither winners nor losers.
worst_day
instance-attribute
¶
Largest single-day LOSS as a signed (negative) figure. This is the number a Daily Loss Limit would have had to absorb.
best_day
instance-attribute
¶
Largest single-day gain, signed. Mirrors BacktestResult.best_day
except that the kernel floors that one at zero for the consistency rule
and this one does not.
p95
instance-attribute
¶
Daily-P&L percentiles, ascending. p05 is the left tail — the bad
day you should size against, as opposed to worst_day, which is the
single realization you happened to draw.
stdev
instance-attribute
¶
Standard deviation of daily P&L, in DOLLARS.
The scalar the percentiles above cannot give you: a Daily Loss Limit is a
fixed dollar threshold, so a dispersion figure in the same units is what
turns "could I breach it" into an answerable question rather than a hope.
A DLL sitting one stdev below the mean day gets hit far more often
than most strategies' authors expect.
POPULATION basis (divides by closed_days, not closed_days - 1),
matching the downside deviation inside SummaryStats.sortino so the two
are read on the same footing. Not annualized, for the reason given on
SummaryStats.calmar. None with no closed days; exactly 0 with
one, which is an artifact of the sample size and not a finding.
best_day_pct_of_profit
instance-attribute
¶
best_day / total_profit as a fraction — the consistency-rule ratio
(docs/topstep-rules.md §4) that must stay at or below
CombineParams.consistency_pct. None when total_profit <= 0:
the rule is only meaningful against a positive profit total, and dividing
by a loss would print a sign-flipped ratio that reads as passing.
RoundTripStats
¶
Bases: Struct
Flat-to-flat trade statistics — the colloquial "per trade" basis.
Everything here is NET of fees, which is the opposite convention to the
half-turn statistics on SummaryStats (those are gross). That is
deliberate, not an oversight: a round trip is a complete decision, so the
honest question is what it earned after costs. A strategy can show a
gross profit_factor above 1 and a negative round-trip expectancy,
and when it does, the round-trip figure is the one that pays you.
Round trips also differ from closed_trades in COUNT: that counts
closing half-turns, so a scale-out of three clips is three closed trades
but one round trip.
count
instance-attribute
¶
Completed flat-to-flat excursions. Positions still open at end of run are excluded — the engine flattens at the session roll, so a finished run normally leaves none.
with_known_risk
instance-attribute
¶
How many of them carried a bracket stop at entry, and therefore have an
R-multiple. The gap between this and count is how much of the R
statistics you should distrust.
avg_loss
instance-attribute
¶
NET-basis win rate and mean win/loss per round trip (avg_loss is a
positive magnitude). None when the relevant side has no trips.
expectancy
instance-attribute
¶
NET dollars earned per round trip — the single number that decides
whether the strategy makes money. None with no trips.
expectancy_r
instance-attribute
¶
Mean R-multiple over round trips WITH a known initial risk.
This is the true R: net P&L over the dollars actually risked at entry,
per excursion. Distinct from SummaryStats.expectancy_r, which is a
gross-basis approximation using R = average loss and exists for strategies
that set no stops. When both are available, prefer this one. None when
no trip had a bracket stop.
worst_r
instance-attribute
¶
Extremes of the R distribution. worst_r below -1 means a stop was
jumped — a gap, or the Tier-0 fill model's stop slippage — and is worth
reading before trusting any risk-based position sizing.
worst_trade
instance-attribute
¶
Extremes of the NET DOLLAR distribution. None with no trips.
Read worst_trade against the account's limits before anything else on
this struct. The R-multiple extremes above say how badly an excursion went
relative to its own plan, which is a question about the strategy; these
say whether the ACCOUNT could absorb it, which is a question about
survival. A single excursion that loses more than the Daily Loss Limit
ends the day by itself no matter how good its R looked, and one that eats
a large share of SummaryStats.distance_to_floor ends the combine.
max_duration_ns
instance-attribute
¶
Mean and longest holding time per excursion, in NANOSECONDS
(closed_ts_ns - opened_ts_ns) — the engine's native stamp, left
unconverted here so no rounding happens before the renderer. Divide by
60_000_000_000 for minutes. None with no trips.
Holding time is a rule question in a combine, not a curiosity: the engine flattens everything at 16:10 ET, so a strategy whose typical excursion approaches the session's remaining length is one that will regularly be closed by the flatten at whatever the tape offers rather than by its own exit. Tier-0 caveat: both endpoints are fill stamps, and a fill is stamped at its bar's OPEN or CLOSE (never in between), so an excursion opened and closed inside one bar reports 0 rather than its true sub-bar life.
DrawdownStats
¶
Bases: Struct
Drawdown under the three reference conventions prop firms actually use.
These are three different numbers on the same price path and a strategy can survive one while violating another, so none of them is "the" drawdown:
staticmeasures against a FIXED initial balance — how far underwater you ever went, ignoring any profit you had banked first.eod_trailingmeasures against a high-water mark that ratchets only on END-OF-DAY closed balance. This is Topstep's actual MLL mechanic.intraday_trailingratchets on intraday equity highs including UNREALIZED profit. This is the Apex-style convention, and it is strictly the harshest of the three.
static
instance-attribute
¶
Deepest point below the STARTING balance: max(0, starting - min
equity). Zero for a run that never traded below its start.
eod_trailing
instance-attribute
¶
Deepest decline from a high-water mark that ratchets only on closed
end-of-day balances (DayRecord.eod_balance, peak seeded at the
starting balance). Intraday excursions are invisible here BY
CONSTRUCTION — that is what makes it the EOD convention.
intraday_trailing
instance-attribute
¶
Deepest decline from a high-water mark that ratchets on intraday
realized+unrealized equity (BacktestResult.bar_equity). None when
that capture is absent (a hand-built BacktestResult, or a run from
before the field existed).
Tier-0 caveat: sampled at the four points of the modelled intrabar path (open, both extremes, close), not from real ticks. The path is deliberately adverse-ordered, so this is a conservative estimate of a real tick-based figure, not a reproduction of one.
eod_episodes
instance-attribute
¶
How many separate EOD drawdown episodes the run contained. An episode
runs from the first closed day below the high-water mark to the day the
mark is regained, or to end of run if it never is. The denominator for
avg_eod_trailing, stated explicitly so a mean over two episodes is not
mistaken for a distribution.
avg_eod_trailing
instance-attribute
¶
Mean depth of those episodes, each measured at ITS OWN trough.
eod_trailing is the maximum over the same set, and the gap between the
two is the shape of the risk rather than its size. A max far above this
mean is one bad week inside an otherwise quiet curve. A max close to it
means the curve simply lives at that depth — the worst case IS the normal
case, and the next episode has no particular reason to be shallower.
None when the run never closed a day below its high-water mark.
longest_days
instance-attribute
¶
Longest unbroken run of closed days spent below the EOD high-water mark. Duration, not depth: a shallow drawdown you sit in for six weeks burns an evaluation window just as effectively as a deep one.
time_to_recovery_days
instance-attribute
¶
Closed days from the DEEPEST EOD trough back to a new high-water mark.
None when that trough never recovered by end of run — which is the
common case and must not be read as "recovered instantly".
min_floor_headroom
instance-attribute
¶
The smallest distance between equity and the trailing MLL floor
reached at any sampled point in the run — how close the account ever came
to termination. None without bar_equity capture.
Distinct from SummaryStats.distance_to_floor, which is the headroom
at the END of the run: a run can finish comfortable having passed within
a tick of the floor mid-way, and only this field shows it.
min_floor_headroom_ts_ns
instance-attribute
¶
ts_init of the bar where min_floor_headroom occurred — the
moment in the evaluation the account was most fragile.
SummaryStats
¶
Bases: Struct
Combine-centric summary metrics for one finished backtest run.
Empty-run conventions: with zero closing half-turns, win_rate,
expectancy and profit_factor are None (undefined, not 0);
profit_factor is also None when there are no losing closes (the
ratio would be infinite). max_drawdown over an empty equity curve
reduces to the terminal mark alone: max(0, starting - ending), which
is 0 for a run with no bars (the balance never moved).
verdict
instance-attribute
¶
verdict: Verdict
Did the strategy pass the evaluation: PASSED, FAILED, or
IN_PROGRESS. Mirrors BacktestResult.verdict so a serialized
SummaryStats carries the outcome it describes.
THREE states, not two. IN_PROGRESS means the tape ran out before
the combine resolved — the strategy neither hit the profit target nor
breached — and it is the OUTCOME OF MOST RUNS. It is not a bad result;
it is an unfinished one. Never collapse this to a boolean by testing
verdict != PASSED: that reports every unfinished run as a failure.
Use :attr:passed / :attr:failed (both False while in progress),
or branch on all three.
closed_trades
instance-attribute
¶
Closing half-turns: trade records with profit_and_loss set and not
voided. NOT round trips — a flip's single half-turn closes one position
and opens the next.
win_rate
instance-attribute
¶
Fraction of closing half-turns with gross profit_and_loss > 0
(fees are charged per half-turn separately, so this is a GROSS stat).
None when there are no closing half-turns.
expectancy
instance-attribute
¶
Mean gross profit_and_loss per closing half-turn. None when
there are no closing half-turns; the aggregate NET counterpart is
net_pnl / closed_trades.
profit_factor
instance-attribute
¶
Sum of gross winning closes / |sum of gross losing closes|. None
when undefined: no closing half-turns, or no losing closes.
max_drawdown
instance-attribute
¶
Largest peak-to-trough decline of the per-bar CLOSE equity curve plus
one terminal mark at ending_balance (the final session roll's flatten
costs land after the last curve point), with the peak seeded at the
starting balance (always >= 0, and never below -net_pnl). Close-basis
only: intrabar excursions are not in BacktestResult.equity_curve.
equity_peak
instance-attribute
¶
Highest equity mark the run reached, close-basis, seeded at the starting balance (so it never reports below it).
Not decoration in a prop account: the peak is what a trailing MLL floor is
ANCHORED to, so this is the number that set the floor you then had to stay
above. Close-basis to match max_drawdown — the two are the opposite
ends of one curve, and equity_peak - max_drawdown is the trough that
produced it.
The real Topstep MLL ratchets on END-OF-DAY closed balances, so the floor
actually in force followed the EOD peak, which is at or below this one.
Where that distinction matters, read drawdown.eod_trailing, which is
measured against that basis.
final_balance
instance-attribute
¶
Ending realized balance (BacktestResult.ending_balance).
net_pnl
instance-attribute
¶
ending_balance - starting_balance, net of ALL fees and commissions
(the engine's session roll flattens at end of run, so nothing is open).
distance_to_floor
instance-attribute
¶
ending_balance - floor: dollars of room above the trailing MLL
floor at end of run.
consistency_headroom
instance-attribute
¶
consistency_pct x total_profit - best_day — dollar slack in the
consistency rule (docs/topstep-rules.md §4). Negative means the best day
is currently too large: the effective target inflates until
best_day <= consistency_pct x total_profit holds.
days_traded
instance-attribute
¶
Closed trading days with trade activity (BacktestResult.days_traded).
exposure
instance-attribute
¶
Fraction of the run's bars during which ANY position was open — a
FRACTION in [0, 1] like win_rate, NOT a percentage. None for a run
with no bars.
This is the figure that tells you how to read every other figure here. Two strategies with identical drawdowns, one at 0.05 exposure and one at 0.95, are not the same risk: the first got that result while off the tape nineteen bars in twenty, and the second has been holding through everything and merely has not met its bad day yet.
A bar counts as exposed when a round trip was open at any point strictly inside it, with the boundary resolved FORWARD — a position opened exactly at a bar's close belongs to the next bar, and an excursion whose open and close carry the same stamp therefore contributes nothing. Multi-instrument runs count each timestamp once: any open contract makes that slice exposed, so this is time-with-risk-on and not a sum of per-symbol exposures. The denominator is every bar the run saw, including bars outside tradable hours.
end_ts_ns
instance-attribute
¶
First and last bar-CLOSE stamp of the run (Bar.ts_init), in epoch
nanoseconds; None for a run with no bars. Provenance — a summary
carrying no window cannot honestly be compared against another one — and
the two ends of :attr:duration_ns.
provisional
instance-attribute
¶
closed_trades < PROVISIONAL_TRADE_FLOOR: this run is too thin for
any distributional metric on it to mean anything.
When True, win_rate, expectancy, profit_factor, payoff_ratio,
sortino, calmar and every daily percentile are still COMPUTED and
still arithmetically correct — they are simply estimates with a standard
error large enough to swamp the effect being measured. The renderer says
so out loud. Treat them as provisional, not as findings.
avg_win
instance-attribute
¶
Mean GROSS P&L of winning closes. None with no winners.
avg_loss
instance-attribute
¶
Mean GROSS loss of losing closes, as a POSITIVE magnitude (so
payoff_ratio is a plain ratio). None with no losers.
payoff_ratio
instance-attribute
¶
avg_win / avg_loss — the size asymmetry that win_rate alone
cannot show. Read the two together, never either alone: a 70% win rate at
a 0.3 payoff ratio is a negative-expectancy strategy waiting for its
sequence. None when either side has no closes.
expectancy_r
instance-attribute
¶
Expectancy expressed in R-multiples, where R is defined as the average losing close — NOT as per-trade initial risk.
This is the approximation available from what the framework records
today. True R-multiples need the stop distance at entry attached to each
round trip; SymbolStrategy.buy/sell accept stop_loss_ticks but do
not retain it, and a strategy that exits on signal has no defined R at
all. Until that lands, read this as "expectancy in units of a typical
loss" — useful for comparing two strategies in this framework,
NOT comparable to an R-multiple quoted anywhere else. None when there
are no losing closes.
longest_losing_streak
instance-attribute
¶
Longest unbroken run of losing closes. A scratch close (exactly 0) is not a loss and BREAKS the streak.
breakeven_cost_per_half_turn
instance-attribute
¶
net_pnl / trade_count: the ADDITIONAL cost per half-turn, on top of
the fees already charged, that would drive this run to exactly zero.
Positive is the slack you have; negative means the run is already
underwater and the figure is how much per half-turn you would have to
SAVE to break even. Per half-turn, not per round trip, because that is
how fees are actually charged (see the module docstring). None with
no half-turns.
sortino
instance-attribute
¶
Mean daily P&L / downside deviation of daily P&L, target 0.
Daily-dollar basis and NOT annualized. Sortino rather than Sharpe
because prop rules punish the downside path specifically and are wholly
indifferent to upside variance. Downside deviation divides by the count
of ALL closed days (the standard convention), not just losing ones.
None with no closed days or no downside deviation (nothing to be
punished for).
calmar
instance-attribute
¶
total_profit / max_drawdown over the run.
Combine-horizon basis and NOT annualized, which is a deliberate
deviation from the conventional annualized-return form: annualizing a
twenty-day sample produces a number with no defensible meaning. Read it
as "profit earned per dollar of worst decline." None when
max_drawdown is zero.
drawdown
instance-attribute
¶
drawdown: DrawdownStats
Drawdown under all three prop-firm conventions — see
:class:DrawdownStats.
round_trips
instance-attribute
¶
round_trips: RoundTripStats
Flat-to-flat trade statistics, NET basis, with true R-multiples where
the entry carried a bracket stop — see :class:RoundTripStats. Note the
basis flip: these are net, the half-turn figures above are gross.
duration_ns
property
¶
end_ts_ns - start_ts_ns: the run's wall-clock span, in
nanoseconds. None for a run with no bars.
CALENDAR time, including every night, weekend and holiday the market
was shut — it says how far apart the run's ends were, not how much
trading it contains. days_traded is the figure to judge a run's
length by, and the two diverge sharply on any sparse feed.
passed
property
¶
The combine was actually cleared. False for IN_PROGRESS —
see :attr:verdict, and do not read not passed as "failed".
failed
property
¶
The combine was actually blown (an MLL/DLL breach, or a rule
violation the kernel treats as terminal). False for
IN_PROGRESS: missing the profit target is not a failure, and a
run that simply ended is neither passed nor failed.
TradeCloseStats
¶
Bases: Struct
Per-closing-half-turn statistics, GROSS basis (see SummaryStats).
Extracted as its own struct so the replay recorder's running snapshots use
the SAME code path compute_summary does — over a prefix of the closes
instead of all of them — and cannot drift from it. Field semantics match
the SummaryStats fields of the same names exactly.
EodDrawdown
¶
Bases: Struct
The EOD-trailing pieces of DrawdownStats, computable on any prefix
of a run's closed days. Extracted so the replay recorder's snapshots share
this fold with compute_summary instead of re-implementing it.
compute_trade_close_stats
¶
compute_trade_close_stats(closes: Sequence[Decimal]) -> TradeCloseStats
Gross per-close statistics over closes (each a closing half-turn's
profit_and_loss). Pure; callable on any prefix of a run's closes.
Source code in src/topstep_backtest/metrics/stats.py
compute_daily_stats
¶
compute_daily_stats(day_records: Sequence[DayRecord], total_profit: Decimal) -> DailyStats
DailyStats over day_records. Pure; callable on any prefix of a
run's closed days (the replay recorder's snapshots do exactly that).
Source code in src/topstep_backtest/metrics/stats.py
compute_eod_drawdown
¶
compute_eod_drawdown(day_records: Sequence[DayRecord], *, starting: Decimal) -> EodDrawdown
EOD-trailing drawdown fold over closed days (see DrawdownStats).
Source code in src/topstep_backtest/metrics/stats.py
compute_round_trip_stats
¶
compute_round_trip_stats(trips: Sequence[RoundTrip]) -> RoundTripStats
Round-trip statistics. NET basis throughout — see :class:RoundTripStats.
Pure; callable on any prefix of a run's completed round trips.
Source code in src/topstep_backtest/metrics/stats.py
exposure_fraction
¶
exposure_fraction(bar_stamps: Sequence[int], trips: Sequence[RoundTrip]) -> Decimal | None
Fraction of distinct bar stamps with a position open.
See :attr:SummaryStats.exposure for the boundary convention. None
when there are no bars — an empty run has no denominator, which is not the
same claim as "was never exposed".
Source code in src/topstep_backtest/metrics/stats.py
sortino_ratio
¶
Mean daily P&L over downside deviation, target 0, NOT annualized.
Pure; callable on any prefix of a run's daily P&Ls.
Source code in src/topstep_backtest/metrics/stats.py
compute_summary
¶
compute_summary(result: BacktestResult, *, trades: Sequence[HalfTradeModel], consistency_pct: Decimal) -> SummaryStats
Derive SummaryStats from a run's frozen result and trade list.
Pure and deterministic: a function of its arguments only (Decimal
arithmetic at the default context), no clock reads, no randomness.
trades is the broker's half-turn list (SimBroker.trades);
consistency_pct comes from the CombineParams the run used.