Skip to content

Quickstart

From an empty environment to a rule-checked verdict on your own data. No prior context assumed. For the same ground covered slowly, with the reasoning behind each decision, read the EMA crossover tutorial instead.

1. Install

pip install topstep-backtest

Python 3.12 or newer. TA-Lib is a core dependency and ships wheels for common platforms; on anything else, install the TA-Lib C library first. Add the data extra to feed in a pandas DataFrame — it carries pyarrow too, so Parquet exports load as well as CSV:

pip install "topstep-backtest[data]"

2. Write a strategy

You subclass SymbolStrategy and implement on_bar. The chassis handles the parts that are easy to get subtly wrong: indicators registered with use() are updated exactly once per matching bar in registration order, and on_bar is gated until every one of them has seen its full lookback, so you cannot read a half-warmed-up average.

from topstep_backtest import SymbolStrategy
from topstep_backtest.indicators import Cross, Sma
from topstep_backtest.protocols import Bar


class SmaCross(SymbolStrategy):
    def __init__(self, contract_id: str) -> None:
        super().__init__(contract_id)
        self.fast = self.use(Sma(10))  # TA-Lib SMA
        self.slow = self.use(Sma(30))  # TA-Lib SMA
        self.cross = self.use(Cross(self.fast, self.slow))  # framework helper

    async def on_bar(self, bar: Bar) -> None:
        if self.cross.up and self.position.flat:
            await self.buy(1, stop_loss_ticks=40, take_profit_ticks=80)
        elif self.cross.down and self.position.is_long:
            await self.close()

stop_loss_ticks and take_profit_ticks attach a signed-tick OCO bracket at entry. Setting them is what gives the report a true R-multiple later — without a stop at entry there is no defined risk to divide by, and the framework reports None rather than inventing one.

Managing the trade after it opens

The bracket is not an attachment to the position — when the entry fills, the venue creates two real working orders, a reduce-only stop and a reduce-only limit, OCO-paired, each with its own id. So they can be amended while the trade runs:

async def on_bar(self, bar: Bar) -> None:
    entry = self.position.avg_price  # the venue's average, None when flat
    if self.position.is_long and entry is not None:
        if bar.close - entry >= 20 * self.spec.tick_size:
            await self.move_stop(ticks=0)  # stop to breakeven

move_stop and move_target take either price= (an absolute level, passed to the broker as given) or ticks= — an offset from the average entry, signed in the position's favour, so ticks=0 is breakeven and ticks=10 is ten ticks of locked profit whether you are long or short. Each returns how many orders it moved: 0 means there was nothing to move, which is what an entry placed without stop_loss_ticks gets. Rejections go to on_reject like every other sugar call, and the remaining orders still move.

They act on bracket children only — orders carrying parent_order_id, the gateway's own shape for "this exists because that entry filled". A stop-entry you placed yourself is deliberately not one of them, which is what stops a breakout strategy's resting order from being mistaken for protection. stop_orders and target_orders expose the same filtered view; anything outside it is working_orders plus ctx.orders.modify(...).

Two things worth knowing. An amended level takes effect from the next bar — the current bar's fill walk ran before on_bar was called, and the accepted_ts firewall is what keeps it that way. And each opening fill gets its own bracket pair, so after a scale-in there are several stops; these methods move all of them, which is why they return a count rather than a single order.

Indicators are TA-Lib, not reimplementations

Sma(10) is TalibIndicator("SMA", timeperiod=10). Nothing in this project implements an indicator formula, so there is no second implementation to drift from the reference. Cross is the deliberate exception — TA-Lib has no crossover primitive, so it compares two TA-Lib outputs rather than computing anything. See Indicators.

Trading one session off a 24h tape

Futures print around the clock, but you may only want to act during one session — and that raises a question the bar loop cannot answer for you: which bars should feed each indicator? There are two independent switches for it.

from topstep_backtest import NEW_YORK, SymbolStrategy
from topstep_backtest.indicators import Atr, Ema


class NyTrend(SymbolStrategy):
    def __init__(self, contract_id: str) -> None:
        super().__init__(contract_id, trade_sessions=(NEW_YORK,))  # when it may TRADE
        self.trend = self.use(Ema(50))  # sees the whole tape
        self.atr = self.use(Atr(14), session=NEW_YORK)  # sees NY bars only

The Ema keeps consuming Asia and London while every decision happens in New York. Restricting trading never starves an indicator — one fed only the tradable window would develop gaps and compute a different value from the same tape.

Which indicators to scope has a rule behind it. Dispersion measures (Atr, StdDev, Rsi) describe how much price moves per bar, and that is session-dependent: on a 24h feed an Atr(14) read at 09:30 ET is computed almost entirely from thin pre-market bars, so it understates New York volatility exactly when you are sizing a stop. Level measures (Sma, Ema) answer where price is, and the overnight move is real. Levels are continuous across sessions; dispersion is not.

Scoping needs 24h data, and costs warmup

data/sample_mnq_1m.csv is RTH-only — 390 bars a day, 09:30–15:59 ET — so it contains no Asia or London bars and nothing scoped is observable on it. For synthetic bars pass synthetic_bars(..., hours="globex").

A scoped indicator also warms in its own cadence: it needs history_bars bars of its session, so on 5-minute bars an NY-scoped Sma(30) spans ~25 trading days against ~7 unscoped. Check strategy.warm, never bars_seen >= history_bars.

ASIA, LONDON and NEW_YORK are defined in their own timezones, so daylight saving comes from the IANA database — London and New York switch on different dates, and the London window really is 04:00 ET rather than 03:00 for about three weeks each spring. Worked end to end in examples/session_scoped.py.

3. Run it on synthetic bars

Synthetic bars are seeded and deterministic, which makes them right for checking that your wiring works before you introduce the confounder of real data.

from datetime import date
from decimal import Decimal

from topstep_backtest import AccountSize, Backtest
from topstep_backtest.core.instruments import spec_for_symbol
from topstep_backtest.data.synthetic import synthetic_bars

MNQ = "CON.F.US.MNQ.U26"

bars = synthetic_bars(
    contract_id=MNQ,
    spec=spec_for_symbol("MNQ"),
    start_day=date(2026, 5, 4),
    days=5,
    seed=7,
    start_price=Decimal("23000.00"),
    bars_per_day=120,
    vol_ticks=12,
)

report = Backtest(bars, SmaCross(MNQ), account=AccountSize.S50K).run()
print(report)

run() is synchronous — it wraps asyncio.run. Inside an existing event loop (a notebook, a service) await bt.arun() instead.

The knobs on Backtest are deliberately few, and they are only things that genuinely exist: the account tier, whether the optional Daily Loss Limit is on, validation strictness, and the fill/broker/fee fidelity configs.

Backtest(bars, SmaCross(MNQ), account=AccountSize.S150K, dll_enabled=True).run()

Economics are never a knob. You cannot ask for cheaper fees or a friendlier fill model to make a strategy pass.

4. Read what comes back

print(report) renders a verdict, the day-by-day trailing-drawdown trail, and four statistics blocks. The single most important thing to know before reading it:

The verdict has THREE states, not two

PASSED, FAILED, and IN_PROGRESS — and IN_PROGRESS is the outcome of most runs. It means the tape ran out before the combine resolved: the strategy neither reached the profit target nor breached a limit. It is an unfinished result, not a bad one. Never read not passed as "failed".

Reading the report walks every block and every number. It is the page to read before you draw a conclusion from any of them.

5. Point it at your own data

bars_from_dataframe (pandas) and bars_from_records (no pandas) turn candles into validated Bar streams:

from topstep_backtest.data.wrangler import bars_from_dataframe

bars = bars_from_dataframe(df, contract_id=MNQ, stamp="close")

Both make you declare stamp="open" or stamp="close". This is not boilerplate. What your timestamps mean is the difference between a causal backtest and one with an off-by-one-bar look-ahead, and no library can infer it from the numbers. Declare it wrong and every result you get is optimistic fiction.

Validation runs by default and raises DataValidationError on defects that would make the run meaningless. See data.validator for which defects are fatal and which are merely worth knowing about.

For quarterly futures, stitch contracts into one continuous series with data.continuous rather than concatenating them — a raw concatenation invents a gain or loss at every roll that no trader experienced.

A Parquet export produced by databento-data-playground's convert.py skips the declarations entirely: the product, bar span and stamp ride in the file's own embedded metadata, and data.loaders reads them from there — after cross-checking the file's tick economics against the engine's, so a mislabeled file is refused rather than silently rescaling every dollar figure:

from topstep_backtest.data.loaders import load_bars

bars, spec, meta = load_bars("data/MNQ.ohlcv-1m.parquet")
report = Backtest(bars, SmaCross(spec.symbol), account=AccountSize.S50K).run()

examples/run_real_data.py is a runnable version of this step, including the --symbol cross-check that catches a contract specification mismatch before it silently rescales every P&L figure in your report. It auto-detects a self-describing export and takes the load_bars path with no flags at all.

Next

  • The workflow — the stages from here to a decision you can defend, and the gate each must pass before the next number means anything.
  • Reading the report — what every number means and which ones mislead when read alone.
  • Beyond one backtest — one run is one sample; here is what to do about that.
  • Topstep rules — what is enforced, with citations and calibration status.
  • Working manual — the dense day-to-day reference, including every metric's basis and the traps in each subsystem.