Catalog

Datasets

Data quality is not a support function for a backtesting platform — it is most of the result. A flawless engine running against a survivorship-biased universe produces a confident, precise, wrong answer. This is the catalog we are assembling, and the state each piece is actually in.

Status: two live, three planned, none purchasable

Two datasets below are real and running: the macro series (published to /macro, regenerated by the engine) and the survivorship-free equity mirror, which is marked internal because it is licensed vendor data we are not permitted to redistribute. The remaining three are planned — sources are named so the intent is checkable, but availability, licensing, and access mechanics are unsettled.

Live does not mean purchasable. There is no account system, no API, and no delivery mechanism of any kind today, so no dataset here is available to anyone outside the operator regardless of its status below.

Survivorship-free US equities

Live (internal)
Source
Sharadar
Coverage
US listed equities 1997-present, point-in-time universe including 15,628 delisted securities
Cadence
Daily

The single most important dataset in the catalog. A point-in-time universe that retains companies which were later acquired, delisted, or went to zero is what separates an equity backtest that means something from one that has quietly pre-selected for survivors.

Includes the delisting and corporate-action history needed to reconstruct the tradable universe as it stood on any past date, rather than as it stands today.

Full-market daily bars

Planned
Source
To be confirmed
Coverage
Open / high / low / close / volume across the full listed universe
Cadence
Daily

The base layer everything else sits on. Breadth matters more than depth here — a strategy screened against a hand-picked subset of tickers is being tested on a universe that was already curated by someone with hindsight.

Split- and dividend-adjusted at the data layer, with the adjustment method stated explicitly. Unadjusted series retained alongside, because which one is correct depends on what is being tested.

Crypto tick archives

Planned
Source
Binance public data
Coverage
Trade-level ticks and order book snapshots for major pairs
Cadence
Continuous

Crypto trades continuously and its microstructure is not the same as an equity session. Tick-level data is what makes a realistic fill model possible for strategies whose edge is measured in basis points.

Venue-specific. A fill model calibrated on one exchange's book does not transfer to another without re-calibration, and results will be reported per venue rather than pooled.

Macro series

Live
Source
FRED (Federal Reserve Bank of St. Louis)
Coverage
18 series across housing, energy, rates, inflation, labor, and monetary aggregates
Cadence
Daily to quarterly, per series

Regime detection needs macro state. Most published macro backtests leak because they use the current revised value of a series rather than the value that was actually printed at the time.

Published now — see the macro page. Every series carries its observation date and a freshness verdict judged against its own release cadence, not a flat clock. Vintage-aware history (ALFRED first prints, rather than today's revised values) is NOT yet implemented: the current artifact serves latest-revision values, which is fine for monitoring and wrong for backtesting a strategy that would have seen the first print. That gap is stated here rather than glossed.

Fundamentals & filings

Planned
Source
SEC EDGAR
Coverage
Company filings and structured financial statement data
Cadence
As filed

Fundamental factors are only testable if the data is keyed to the filing date rather than the fiscal period it describes. A quarter ending in March is not knowable in March.

Indexed by filing timestamp, not period end, specifically to prevent the most common look-ahead error in fundamental backtesting.

Principles

Three rules govern anything that enters the catalog, and they are the same rules that make the backtesting doctrine enforceable rather than aspirational.

Access

Dataset access is included from the Pro tier upward, delivered through the engine API, with each set becoming available as it is released. Neither the API nor any purchasable tier exists yet — see the waitlist to be told when it does. We will not redistribute third-party data in violation of the licence it came under, so some sources may end up available only as derived features rather than as raw series.