Reading the report¶
print(report) renders everything one run knows about itself. This page walks it top to
bottom, says what basis each number is on, and names the ones that mislead when read alone.
If you only read one section, read what is deliberately absent — the metrics a conventional backtester prints and this one refuses to, and why each refusal matters more in a funded-account evaluation than it would anywhere else.
The whole thing¶
This is real output from examples/ema_cross.py, not an illustration:
== Topstep Combine 50K: IN_PROGRESS ==
combine still in progress at end of data
window 2026-05-04 09:31 ET -> 2026-05-11 11:30 ET (7d 01:59:00 calendar)
balance 50000.00 -> 49995.16 (net -4.84)
total profit -4.84 (target 3000.00)
best day 84.04 (cap -2.42 = 50% of total)
MLL floor 48000.00 (distance 1995.16)
days traded 5 closing trades 8 half-turns 16
warmup 26 bars gated before the first decision
day trail
2026-05-04 eod= 49981.04 pnl= -18.96 floor= 48000.00 *
2026-05-05 eod= 49972.56 pnl= -8.48 floor= 48000.00 *
2026-05-06 eod= 49955.08 pnl= -17.48 floor= 48000.00 *
2026-05-07 eod= 49911.12 pnl= -43.96 floor= 48000.00 *
2026-05-08 eod= 49995.16 pnl= +84.04 floor= 48000.00 *
2026-05-11 eod= 49995.16 pnl= +0.00 floor= 48000.00
summary stats
verdict IN_PROGRESS
closed trades 8
!! PROVISIONAL — 8 closes is under the 200 this framework requires before
treating a measured edge as distinguishable from sampling noise. Every
rate, ratio and percentile below is an ESTIMATE, not a finding.
win rate (gross) 37.50%
expectancy (gross) +4.38
expectancy (R=avg L) 0.37
avg win / avg loss 31.33 / 11.80
payoff ratio 2.66
profit factor (gross) 1.59
longest losing streak 5
net P&L -4.84
breakeven cost/half -0.30
sortino (daily) -0.04
calmar (profit/DD) -0.04
distance to floor 1995.16
consistency headroom -86.46
days traded 5
equity peak 50032.76 <- what a trailing floor anchors to
exposure 28.33% of bars held a position
drawdown (three conventions — a strategy can survive one and violate another)
static (from start) 105.12
EOD trailing 88.88 <- Topstep's MLL mechanic
EOD trailing (avg) 88.88 <- mean of 1 episode(s); near the max = the norm
intraday trailing 147.88 <- harshest; modelled path, not ticks
peak-to-trough 137.88 <- close-sampled; the legacy `max_drawdown`
longest 6 day(s) below the EOD high-water mark
time to recovery not recovered
min floor headroom 1890.88 (2026-05-08)
round trips (8 flat-to-flat — NET of fees, unlike the gross figures above)
win rate (net) 37.50%
expectancy (net) -0.60
avg win / avg loss 28.85 / 18.28
payoff ratio 1.58
expectancy (true R) -0.02 (8/8 with a stop at entry)
best / worst R 1.94 / -0.94
best / worst trade +77.52 / -37.48 <- read the worst against the DLL
holding time avg/max 01:00:22 / 04:45:00
daily P&L (6 closed: 1 win / 4 lose / 1 flat)
worst -43.96
p05 -43.96
p25 -18.96
median -17.48
p75 +0.00
p95 +84.04
best +84.04 (consistency ratio undefined: no net profit to share)
stdev 40.27 <- dollars, so it compares directly against the DLL
topstep-backtest 0.2.0 — unofficial simulation, not affiliated with Topstep. Rule and fee
constants are cited config, NOT calibrated against a live account: treat the verdict as a
diagnostic, not an authoritative pass/fail.
Note what this run demonstrates: a gross profit factor of 1.59 that still loses money. That is not a contradiction, it is the basis rules doing their job. Read on.
Three rules that govern every number¶
Most trading metrics admit two honest answers, and mixing them silently is how a report lies. Three distinctions run through this entire output.
Gross versus net. Fees and commissions are charged on every half-turn — entry and
exit both — and deducted from the balance separately from realized P&L. So the summary
stats block is gross: its win rate, expectancy, payoff ratio and profit factor are
computed before costs. The round trips block is net. net P&L is the only net figure
in the gross block. This is exactly why the run above shows profit factor 1.59 and
net P&L -4.84.
Half-turns versus round trips. A "closing half-turn" is one broker fill that realized P&L. A "round trip" is a complete flat-to-flat excursion. They differ in count as well as basis: a three-clip scale-out is three closed trades but one round trip. When somebody says "trade", they usually mean a round trip; when a broker says it, they mean a half-turn.
Close-sampled versus intrabar. Tier-0 fidelity means bars, not ticks. Figures derived from the bar-close equity curve are exact for what they measure; figures labelled intrabar are sampled along one modelled pessimistic path (open → adverse extreme → other extreme → close), which makes them conservative estimates rather than reproductions of a real tape.
The header¶
verdict — PASSED, FAILED, or IN_PROGRESS. Three states, and the third is the
outcome of most runs: the tape ran out before the combine resolved. It is an unfinished
result, not a bad one. report.stats.passed and .failed are both False there, so
never collapse this to a boolean by testing not passed.
window — first to last bar close, in Eastern time, with the calendar span. Calendar,
including every night and weekend the market was shut; judge a run's length by days traded
instead. The window exists so a saved summary carries the period it describes — a metric
without its window cannot honestly be compared to another one.
balance / total profit — realized balance start to end, net of all fees. total profit
is measured off the last end-of-day close, against the target for the account tier.
best day / cap — the consistency rule: your single best day must be no more than
50% of total profit. When total profit is negative the cap is negative too and the rule is
nowhere near satisfied, which is what the example shows. consistency headroom further down
is the dollar slack; negative headroom means a balance that reached the target still does
not pass.
MLL floor / distance — the trailing Maximum Loss Limit floor and how far above it you ended.
days traded / closing trades / half-turns — 16 half-turns is 8 entries plus 8 exits. Fees are charged per half-turn, so that is the number that costs you money.
warmup — bars gated before the strategy's first decision. If this equals your bar count, your indicators need more history than you supplied and the strategy never ran.
A REJECTED line is a wiring bug, and its absence is information
When the broker refuses an order placement, the report grows a REJECTED line counting
refusals per gateway error code. Zero trades with rejections is broken wiring; zero
trades without is a quiet strategy. They look identical otherwise. Common codes: 4
position cap or dead account, 5 the 16:10–18:00 ET no-trade window or a weekend, 8
unknown contract, 2 validation.
The day trail¶
Per closed trading day: end-of-day balance, that day's P&L net of fees, the MLL floor after
that day's ratchet, and * if the day had trades.
Watch the floor column. It moves only on a new equity high at end of day — Topstep's MLL is two-state and ratchets on closed balances, then locks permanently once you are far enough ahead. Intraday-trailing is Apex's convention, not Topstep's; do not port that intuition. In the example the floor never moves, because the run never made a new EOD high.
Summary stats (gross)¶
| Figure | What it is | The trap |
|---|---|---|
win rate |
Fraction of closing half-turns that were profitable | Meaningless alone — read it with payoff ratio |
expectancy |
Mean gross P&L per closing half-turn | Gross; the net counterpart is in the round-trips block |
expectancy (R=avg L) |
Expectancy in units of the average loss | Not a real R-multiple; prefer the round-trip true R |
payoff ratio |
avg win / avg loss |
A 70% win rate at 0.3 payoff is negative expectancy waiting for its sequence |
profit factor |
Gross wins ÷ gross losses | Gross — can sit above 1 on a losing run |
longest losing streak |
Longest unbroken run of losing closes | A scratch (exactly 0) breaks the streak |
breakeven cost/half |
Extra per-half-turn cost that would zero the run | Negative means you are already under water by that much |
sortino |
Mean daily P&L ÷ downside deviation | Daily-dollar basis, not annualized |
calmar |
Total profit ÷ max drawdown | Combine-horizon basis, not annualized |
equity peak |
Highest close-basis equity mark | What a trailing floor anchors to; the real ratchet is EOD |
exposure |
Fraction of bars that held a position | A fraction, not a percent, in the API |
exposure is the one to read first, because it tells you what everything else is a
sample of. Two strategies with identical drawdowns, one at 0.05 exposure and one at 0.95,
are not the same risk: the first got that result while off the tape nineteen bars in twenty,
and the second has been holding through everything and has merely not met its bad day yet.
PROVISIONAL is not a formatting flourish
Below 200 closing half-turns the report prints a provisional banner. Nothing is suppressed and every figure remains arithmetically correct — they are estimates whose standard error swamps the effect being measured. The example run has 8. Do not quote a provisional metric as a finding.
Drawdown¶
Three conventions, deliberately, because they are different numbers on the same price path and a strategy can survive one while violating another:
static— depth below the fixed initial balance. Ignores profit you banked first.EOD trailing— depth below a high-water mark that ratchets only on closed end-of-day balances. This is Topstep's actual MLL mechanic.intraday trailing— ratchets on intraday equity including unrealized profit. The Apex-style convention, and always the harshest of the three.peak-to-trough— a fourth, close-sampled measure; the legacymax_drawdownfield.
In the example these are 105.12 / 88.88 / 147.88 / 137.88 on one run. Always say which one you mean.
EOD trailing (avg) averages episodes — each separate excursion below the high-water
mark, measured at its own trough — while EOD trailing is the maximum of that same set. The
gap between them is the shape of the risk rather than its size: a max far above the mean is
one bad week inside an otherwise quiet curve; a max close to the mean means the curve lives
at that depth and the next episode has no particular reason to be shallower.
min floor headroom is how close the account ever came to termination, as against
distance to floor, which is only where it ended. A run can finish comfortable having
passed within a tick of death mid-way, and only this field shows it.
Round trips (net)¶
The basis flips here deliberately: a flat-to-flat excursion is a complete decision, so the
honest question is what it earned after costs. Compare the example's
expectancy (gross) +4.38 per half-turn against expectancy (net) -0.60 per round trip.
The net one is the one that pays you.
expectancy (true R) is net P&L over the dollars actually risked at entry, taken from the
bracket stop. The 8/8 with a stop at entry annotation is the number to check first — R
statistics computed from a minority of trips are not describing your strategy. A
worst R below −1 means a stop was jumped, by a gap or by the fill model's stop slippage.
best / worst trade is the same excursions in dollars, and it answers a different
question. R asks how a trade went against its own plan; dollars ask whether the account
could absorb it. A −0.94R loss sounds disciplined until you notice it was −37.48 against a
daily loss limit. Check the worst trade against the DLL and against distance to floor
before you trust either ratio.
holding time matters because the engine flattens everything at 16:10 ET. A strategy whose
typical hold approaches the session's remainder is one the flatten keeps closing at whatever
the tape offers, rather than one exiting on its own signal.
Daily P&L¶
The distribution you need to reason about a Daily Loss Limit. worst is the one bad day you
happened to draw; p05 is the one you should size against. stdev is the same
distribution's spread in dollars, which is what makes it directly comparable to a
dollar-denominated DLL — a limit sitting within one standard deviation of the mean day will
be getting hit routinely.
What this report does not tell you¶
A conventional backtester prints a dozen metrics that are absent here. Each omission is a decision, and in a funded-account context each one is load-bearing.
No Return [%], Return (Ann.) or CAGR. All three divide by a starting balance you
never posted. A combine account is not capital at risk — the evaluation fee is. A percentage
return on $50,000 is a fiction that gets worse the more seriously it is taken. The
economics module provides the honest version: expected
value per attempt, against your own fee and payout assumptions.
No Buy & Hold Return, Alpha or Beta. They presume a benchmark you could actually
have held. Holding ES through a combine breaches the trailing drawdown long before the
window closes, so it is not an alternative that was ever available to you.
No annualized volatility. Annualizing a twenty-day sample produces a number with no
defensible meaning. daily.stdev reports the same underlying quantity in dollars, on the
horizon that was actually measured.
No headline Sharpe ratio. Sharpe exists in this codebase where it does real work — inside deflated Sharpe, corrected for the number of trials you ran. Promoting it to the summary would invite the annualization confusion above, and Sortino is the better headline anyway: prop rules punish the downside path specifically and are wholly indifferent to upside variance.
No percentage-based trade statistics. Conventional backtesters compute best/worst/average
trade, expectancy and profit factor from ReturnPct, the price-relative return against entry
notional. An MNQ trade from 23000 to 23010 is 0.04%, which tells you nothing about the
dollars that moved you toward or away from a fixed threshold. Every trade figure here is in
dollars or in R.
No Kelly criterion. Kelly optimizes long-run growth under continuous fractional sizing and an unbounded horizon. A combine gives you integer contracts, a fixed floor, and a first-passage objective — reach the target before touching the floor. That is a barrier-hitting problem, not a growth problem, and the growth-optimal fraction can be flatly wrong for it. Monte Carlo answers the question Kelly is gesturing at, directly.
No SQN. It is a t-statistic on trade P&L asking whether an edge is distinguishable from noise. Deflated Sharpe and PBO answer that same question with actual multiple-testing correction, and putting a weaker answer beside a stronger one helps nobody.
When the answer is wrong¶
Every item here is a bug that runs clean and returns a plausible number. These are the conditions under which a report on this page is silently wrong rather than merely imprecise.
Your stamp was wrong. bars_from_dataframe(..., stamp=...) declares what your source
timestamps mean. Declaring close for open-stamped candles is a one-bar look-ahead that
inflates every figure in the report and raises nothing. Confirm your vendor's convention
before you pass it — no library can infer it from the numbers.
Your edge lives inside the bar. Tier-0 resolves fills along one modelled path. A bar holding both your stop and your target resolves adverse-extreme-first: the stop wins. If your strategy's profitability depends on which of the two was really touched first, this framework cannot evaluate it, and the number it returns is not a conservative estimate — it is unrelated to the answer.
Your tape contains exchange holidays. There is deliberately no holiday calendar — a
hand-maintained one was wrong on roughly five dates a year and silently deleted tradable
sessions. A holiday bar is therefore indistinguishable from any weekday bar, and
synthetic_bars(days=N) counts weekdays. Filter exchange holidays upstream, in the data you
feed in.
You concatenated quarterly contracts. Splicing raw contract series invents a gain or
loss at every roll that no trader experienced. Use
data.continuous.
You are treating the verdict as authoritative. The MLL breach number is framework-computed and cannot be reconciled against the live gateway: the real SDK exposes no equity or unrealized-P&L field to diff it against. The value deciding pass/fail is computed here and nowhere else, from constants that are cited configuration rather than calibration. This is the single largest parity hazard in the project. Report every verdict as a diagnostic.
You are reading one sample. Even with none of the above, one backtest is one path through one tape. Beyond one backtest is the next page for a reason.
Getting at the numbers¶
Report holds four things, all typed and all serializable:
report.result # the frozen BacktestResult — msgspec.json.encode() it and diff two runs
report.stats # SummaryStats: everything on this page, as Decimals
report.trades # the broker's half-turn list
report.bars_gated # warmup-gated bars, or None for a non-SymbolStrategy
s = report.stats
s.verdict, s.passed, s.failed # three states — see the header section
s.exposure # fraction in [0, 1], or None with no bars
s.drawdown.eod_trailing # Topstep's MLL mechanic
s.drawdown.min_floor_headroom # closest the account came to termination
s.round_trips.worst_trade # net dollars — check against the DLL
s.daily.p05 # the bad day to size against
s.consistency_headroom # negative = a passing balance still fails
Full field-by-field documentation, including every basis and edge case, is in the
metrics.stats reference — those docstrings are the
authoritative version of this page.
The same report, as a chart¶
Everything above also renders as one self-contained HTML file — no server, no network, so it opens offline and archives next to the run:
report.to_html("tearsheet.html") # name the file
report.show() # or write a temp file and open a browser
The sheet adds what text cannot show: the candlestick tape with every fill marked per round trip, the equity curve against the trailing MLL floor with the modelled intrabar envelope, daily P&L as bars, and the R-multiple distribution. Every statistic on this page appears there too, with the same basis labels — both renders share their formatting helpers, so they cannot drift apart on what a number means.
For a sheet from every run without inventing a filename each time:
That writes runs/tearsheet-<UTC stamp>.html. The clock is read for the filename only — the
document itself is the same byte-identical render of frozen run data, and two runs landing in
the same second get -2, -3, … rather than one overwriting the other.
Details and the panel-by-panel breakdown are in the tearsheet
reference.
Replaying a run bar by bar¶
The report and the tearsheet above answer what happened. When the question becomes why — what did the strategy see on that bar, and what did it do about it — record the run:
record=True changes nothing about the run — the report is byte-identical either way,
pinned by a golden test — but the tearsheet now opens with two tabs: results (the sheet
above, the finished run) and replay · bar by bar. The replay tab is a cockpit sized to
the window — charts down one side, the settled state and running stats beside them, the
event log across the full width beneath — so stepping the run never means scrolling between
the tape and the panel explaining it. Step it bar by bar (buttons, slider, arrow keys; space autoplays, at one bar a second
when you want to watch it happen, up to 40 to cover ground) and
it shows only what the strategy had seen: everything after the cursor is veiled on every one
of its charts, while the results tab keeps showing the run that finished. At each bar you
get:
- the live trade numbers on the tape itself — a strip across the top of the candles with
the cursor's bar, time and OHLC, the position and its average entry, open P&L, the day's
P&L, bar-close equity, floor headroom, and the closed-trade record so far (net P&L, net
win rate, trade count) from the running-stats snapshot in force. The
statsbutton in the toolbar hides it; - a viewport that follows the cursor — the second dropdown picks how much tape stays in view (whole run, or a 60–480 bar window; a recording over 400 frames opens following, because a whole run squeezed into one pane gives each bar half a pixel). The price axis carries a line at the cursor's own close, and the equity axis one at the cursor's equity: the series' usual last-value badges are switched off here, since a badge reporting the end of a run the cursor has not reached is future information printed on the axis;
- the state after the bar settled — net position and its average price, working orders (also drawn as price lines on the candlestick pane, stops red, targets green), realized balance, bar-close equity, floor headroom, and the day's running P&L;
- indicator values, named by the attributes your strategy stores them under
(
self.fast = self.use(Sma(20))shows asfast), on their own synced chart, withCrossfires marked on the tape; warmup-gated bars say so instead of pretending to be holes; - running statistics — the last snapshot at or before the cursor, one per closing trade
and day close, split into by-type cards (status · trades (gross) · round trips (net)
· risk & quant · daily P&L) shown one at a time behind filter chips, so the panel
holds one readable card instead of fifty rows. Every row the newest snapshot moved is
marked
•(hover it for the previous value) — a chip whose hidden card just moved carries the same dot — and the header names which snapshot of how many is in force. Each snapshot is the full stats block computed over the run's prefix by the same code that computes the final report, so a scrubbed number can never disagree with the one printed at the bottom of the page. (exposurerefreshes at day closes only — it is the one figure with no exact incremental form.) - the event log, running the full width under the charts — every decision with its exact parameters and the broker's answer (a rejection stays loud: it is the "intent diverged from execution" moment), every fill with gross P&L and costs, every position snapshot, and the enforcement between bars: the 16:10 flatten, the session close, an MLL breach or DLL lockout. Click a row to jump to it, and use the chips beside the heading to drop a kind you are not reading — an order lifecycle is four rows per trade and buries the two that carry the reasoning. Everything starts shown and the header says how many rows are hidden, because a log quietly missing a rejection is worse than a noisy one.
The recorder captures what on its own; only the strategy knows why. That is what
self.note(...) is for — say it at the decision point and read it back at the same bar:
async def on_bar(self, bar):
if self.cross.up and self.position.flat:
self.note(f"fast {self.fast.value} crossed above slow {self.slow.value}: long 2")
await self.buy(2, stop_loss_ticks=40, take_profit_ticks=80)
Notes are a pure sink — a no-op when the run is not recorded, and never able to influence a recorded one.
Two practicalities. Recording holds per-bar values in memory and in the file, so record the
runs you intend to step through, not every sweep trial. And on a long tape the sheet embeds a
loudly labelled window rather than every frame (centred on the MLL breach when there is
one, else the tail); report.to_html(path, replay=(start, end)) picks the window,
replay="full" forces everything, replay="off" renders the classic sheet, and
report.replay_json(path) dumps the raw recording when a browser page is the wrong tool.