Skip to content

data.loaders

Load a databento-data-playground Parquet export whose product, bar span and timestamp semantics ride in the file's own metadata.

loaders

Load a databento-data-playground Parquet export into engine-ready bars.

The file is self-describing: the product, the bar span and which end of the interval ts_event marks all live in its Parquet key-value metadata. They are read from there rather than passed in, because the engine requires all of them with no defaults and each is a silent corruption when guessed — the wrong product scales P&L (an ES file run as MNQ is exactly 25x wrong), the wrong span mis-stamps every bar against the session clock, and the wrong stamp shifts every signal by one bar.

pandas and pyarrow are imported lazily — only load_bars needs them, and both arrive with the [data] extra.

load_bars

load_bars(path: str | Path) -> tuple[tuple[Bar, ...], InstrumentSpec, dict[str, Any]]

Read the export -> (bars, spec, metadata), ready for Backtest(...).

bars, spec, meta = load_bars("data/MNQ.ohlcv-1m.parquet")
Backtest(bars, MyStrategy(spec.symbol), account=AccountSize.S50K).run()

The returned contract id is the bare root ("MNQ"), matching what stitch_continuous labels a continuous series with. meta is the file's full embedded metadata blob (roll days, adjustment note, provenance, ...), returned so callers can log or display it.

Source code in src/topstep_backtest/data/loaders.py
def load_bars(path: str | Path) -> tuple[tuple[Bar, ...], InstrumentSpec, dict[str, Any]]:
    """Read the export -> (bars, spec, metadata), ready for ``Backtest(...)``.

        bars, spec, meta = load_bars("data/MNQ.ohlcv-1m.parquet")
        Backtest(bars, MyStrategy(spec.symbol), account=AccountSize.S50K).run()

    The returned contract id is the bare root ("MNQ"), matching what
    ``stitch_continuous`` labels a continuous series with. ``meta`` is the
    file's full embedded metadata blob (roll days, adjustment note,
    provenance, ...), returned so callers can log or display it.
    """
    pd, pq = _import_pandas_and_parquet()

    path = Path(path)
    metadata: dict[bytes, bytes] = pq.read_schema(path).metadata or {}
    if _METADATA_KEY not in metadata:
        raise ValueError(
            f"{path.name} carries no {_METADATA_KEY.decode()!r} metadata — it was not "
            "produced by databento-data-playground's convert.py"
        )
    meta: dict[str, Any] = json.loads(metadata[_METADATA_KEY])

    try:
        ours: dict[str, Any] = meta["instrument"]
        root: str = ours["root"]
        bar_schema: str = meta["bar_schema"]
        stamp: Any = meta["stamp"]
    except KeyError as exc:
        raise ValueError(
            f"{path.name}: embedded metadata is missing required field {exc.args[0]!r}"
        ) from None

    try:
        spec = spec_for_symbol(root)
    except KeyError:
        raise ValueError(
            f"{path.name}: no built-in InstrumentSpec for product root {root!r}"
        ) from None

    # The engine's spec is the runtime source of truth; the file's is a copy.
    # If they diverge every dollar figure is wrong, so fail rather than pick one.
    if (float(spec.tick_size), float(spec.point_value)) != (ours["tick_size"], ours["point_value"]):
        raise ValueError(
            f"instrument spec mismatch for {root}: file says tick={ours['tick_size']} "
            f"point_value={ours['point_value']}, engine says tick={spec.tick_size} "
            f"point_value={spec.point_value}"
        )

    unit, unit_number = _bar_span(bar_schema)
    bars = bars_from_dataframe(
        pd.read_parquet(path),
        contract_id=root,
        spec=spec,
        unit=unit,
        unit_number=unit_number,
        stamp=stamp,
    )
    return bars, spec, meta