Polymarket Orderbook Archive archive.pendulumflow.com

What is in these files

PrefixFormatOverview
v3/One file per hour plus a manifest V3 data overview
contrib/devpyle/Contributed, 68 hours, four files per hour in the contributor's schema Contributed tier overview, README
third-party/ag6/One file per hour, sixteen typed columns, the same schema as pmxt/v2/ AG6 data overview
pmxt/v2/One file per hour, sixteen typed columns V2 data overview
pmxt/v1/One file per hour, five columns with JSON inside V1 data overview

The differences that matter

pmxt/v1, pmxt/v2v3/
Files per hour11, plus a manifest and, on recent hours, a witness_stats.json timing file
Event orderingMillisecond timestamps that can tie; export order not guaranteed Sort by timestamp_received; sequence is not a global order
De-duplicate byTransaction hash plus fill fields sequence, unique, measured
Receive-time resolutionMillisecondsMicroseconds
Exchange-time resolutionMillisecondsMilliseconds
Per-file checksumsSHA256SUMS.txtEach hour's manifest, and SHA256SUMS.txt
Coverage attestedNoPer hour, see the audit

v1 and v2 overlap between 2026-04-13 and 04-16, and 59 filenames exist in both with different contents. Keep them apart by prefix, never by filename.

The practical consequences are on Working with the historical data.

Downloading from a script

Scripts and other programmatic readers should fetch files from https://dl.pendulumflow.com: v3/, pmxt/v1/, pmxt/v2/ and third-party/ag6/ are served there at the same paths as on this host. Since 3 October 2026 00:00 UTC, parquet URLs under /v3/, /pmxt/v1/ and /pmxt/v2/ on this host redirect to the same path on https://dl.pendulumflow.com. DuckDB, pyarrow and requests follow redirects; curl needs -L. The redirect covers parquet files only: checksum files, manifests, listings and these pages are still served here.

The archive is served from object storage through a CDN. Under heavy load a request may receive 503 with Retry-After: back off and retry. Remote-parquet readers that scan many hours should prefer full-hour files to millions of byte-range reads.

Reading the files faster

Three DuckDB settings made one reader's backtests over these files more than 50% faster. Each helps in one situation and hurts in another, and results depend on the machine: benchmark your own workload with and without them.

SET threads TO 1;
SET enable_external_file_cache = false;
SET memory_limit = '4GB';
SELECT count(*) FROM read_parquet('v3/2026-09-*/*/*.parquet');

Tips shared on our Discord by @david_82262 (threads and the file cache) and @kuinox (the memory budget).

The two grades

The index on the front page sorts these corpora into two, and the difference is what you can conclude from an hour rather than how much of it there is.

Military Grade

Good enough to replay the orderbook. Several machines recorded independently and the merge keeps the union of what they heard, so a gap in one is covered by another. Every row carries a sequence that is unique within its hour, so duplicates collapse exactly rather than approximately. Order by timestamp_received, which is recorded in microseconds: fine enough that events rarely share a timestamp, where a millisecond clock ties constantly. It does not reconstruct the exact order one socket saw - sequence is collector-local provenance, not a global order. Each hour's event count is checked by a separate pass and given a published coverage verdict. Hours not yet checked are listed as such.

Snapshot Grade

One file per hour from a single source. You can see what the book looked like, and for most questions that is enough. What you cannot do is replay it exactly: there is no sequence column, receive times are milliseconds and can tie, so events inside the same millisecond have no recoverable order. No machine corroborates another, and no per-hour coverage verdict has been published for these, so a thin hour looks like a quiet one.

For prices, spreads and depth over time, Snapshot Grade is enough. It is not a replay of the orderbook. In either grade the files show what the book displayed and what printed, not what an order of yours would have received; a backtest that fills the moment a price is touched overstates its fills, see A fill, from a price you touched.

Where each era came from

Five corpora sit end to end, and they were not gathered the same way. Who recorded an hour, and what has been checked since, matters more than its dates.

EraProvenanceSpan
pmxt/v1/Mirrored from PMXT, byte-identical 2026-02-21T18 → 2026-04-16T05
pmxt/v2/Mirrored from PMXT, byte-identical 2026-04-13T19 → 2026-08-09T23
third-party/ag6/One collector, no second machine to cross-check it. Mirrored 2026-08-26 from polymarket-archive.ag6.ai 2026-08-01T13 → 2026-08-15T09
contrib/devpyle/Contributed; not witnessed by this archive. Recorded by devpyle, crypto Up-or-Down markets only, in his own schema. What that means 2026-08-15T10 → 2026-08-18T05
v3/Our own capture, several machines, merged and checked hour by hour 2026-08-18T06 → ongoing

The four eras do not join up. AG6's contiguous run of 134 hours starts at 2026-08-09T20, four hours before PMXT's last hour rather than after it, so those two overlap; its listing begins earlier still, at 2026-08-01T13, with two hours that sit inside PMXT's span. After AG6 ends there is a gap of 68 hours, 2026-08-15T10 to 2026-08-18T05, that nothing here covers for the whole venue - the contributed tier holds only a handful of crypto markets for those hours - and then our own capture begins.

Credit the collector: pendulumflow for V3, PMXT for V1 and V2, AG6 for their V2 archive. Ours and PMXT's are CC BY 4.0; AG6 states no licence. How to cite.

Serving these bytes is not endorsing them. We are not affiliated with Polymarket.

For AI readers: llms.txt, what this archive holds and the questions it cannot answer.

https://x.com/PendulumFlow

Join the Pendulum Flow Discord: other people who build on prediction-market data, comparing notes and sharing tooling and findings.