What is in these files
| Prefix | Format | Overview |
|---|---|---|
v3/ | One file per hour plus a manifest | V3 data overview |
contrib/devpyle/ | Contributed, 68 hours, four files per hour in the contributor's schema | Contributed tier overview, README |
third-party/ag6/ | One file per hour, sixteen typed columns,
the same schema as pmxt/v2/ |
AG6 data overview |
pmxt/v2/ | One file per hour, sixteen typed columns | V2 data overview |
pmxt/v1/ | One file per hour, five columns with JSON inside | V1 data overview |
The differences that matter
pmxt/v1, pmxt/v2 | v3/ | |
|---|---|---|
| Files per hour | 1 | 1, plus a manifest and, on recent hours, a witness_stats.json timing file |
| Event ordering | Millisecond timestamps that can tie; export order not guaranteed | Sort by timestamp_received; sequence is not a global order |
| De-duplicate by | Transaction hash plus fill fields | sequence, unique, measured |
| Receive-time resolution | Milliseconds | Microseconds |
| Exchange-time resolution | Milliseconds | Milliseconds |
| Per-file checksums | SHA256SUMS.txt | Each hour's manifest, and SHA256SUMS.txt |
| Coverage attested | No | Per hour, see the audit |
v1 and v2 overlap between 2026-04-13 and 04-16, and 59 filenames exist in both with different contents. Keep them apart by prefix, never by filename.
The practical consequences are on Working with the historical data.
Downloading from a script
Scripts and other programmatic readers should fetch files from
https://dl.pendulumflow.com: v3/, pmxt/v1/, pmxt/v2/ and
third-party/ag6/ are served there at the same paths as on this host. Since 3 October 2026 00:00 UTC, parquet URLs under
/v3/, /pmxt/v1/ and /pmxt/v2/ on this host redirect to the
same path on https://dl.pendulumflow.com. DuckDB, pyarrow and requests follow redirects; curl needs
-L. The redirect covers parquet files only: checksum files, manifests, listings and
these pages are still served here.
The archive is served from object storage through a CDN. Under heavy load a request may
receive 503 with Retry-After: back off and retry. Remote-parquet readers
that scan many hours should prefer full-hour files to millions of byte-range reads.
Reading the files faster
Three DuckDB settings made one reader's backtests over these files more than 50% faster. Each helps in one situation and hurts in another, and results depend on the machine: benchmark your own workload with and without them.
SET threads TO 1;helps when your own code already keeps every core busy, for example a backtester running many workers in parallel: DuckDB uses all cores by default, and the two compete for them. It can slow a reader that would otherwise benefit from DuckDB's own parallelism.SET enable_external_file_cache = false;helps a pass that reads each file once: from DuckDB 1.3.0, by default, DuckDB caches data it reads from external files such as Parquet in memory, subject to its memory limit (80% of RAM by default), and in a parallel run that can push the machine into swap. It hurts a reader that reads the same file again, which then fetches it again.SET memory_limit = '4GB';gives each DuckDB process a budget with headroom; it does not by itself guarantee that the total of several processes stays below physical RAM. Set it too low and large sorts and joins spill to disk or fail.
SET threads TO 1;
SET enable_external_file_cache = false;
SET memory_limit = '4GB';
SELECT count(*) FROM read_parquet('v3/2026-09-*/*/*.parquet');
Tips shared on our Discord by @david_82262 (threads and the file cache) and @kuinox (the memory budget).
The two grades
The index on the front page sorts these corpora into two, and the difference is what you can conclude from an hour rather than how much of it there is.
Military Grade
Good enough to replay the orderbook. Several machines recorded
independently and the merge keeps the union of what
they heard, so a gap in one is covered by another. Every row carries a sequence
that is unique within its hour, so duplicates collapse exactly rather than approximately.
Order by timestamp_received, which is recorded in microseconds: fine enough
that events rarely share a timestamp, where a millisecond clock ties constantly.
It does not reconstruct the exact order one socket saw -
sequence is collector-local provenance, not a global order. Each hour's event
count is checked by a separate pass and given a published coverage verdict. Hours not yet
checked are listed as such.
Snapshot Grade
One file per hour from a single source. You can see what the book looked like,
and for most questions that is enough. What you cannot do is replay it exactly: there is no
sequence column, receive times are milliseconds and can tie, so events inside
the same millisecond have no recoverable order. No machine corroborates another, and no
per-hour coverage verdict has been published for these, so a thin hour looks like a quiet
one.
For prices, spreads and depth over time, Snapshot Grade is enough. It is not a replay of the orderbook. In either grade the files show what the book displayed and what printed, not what an order of yours would have received; a backtest that fills the moment a price is touched overstates its fills, see A fill, from a price you touched.
Where each era came from
Five corpora sit end to end, and they were not gathered the same way. Who recorded an hour, and what has been checked since, matters more than its dates.
| Era | Provenance | Span |
|---|---|---|
pmxt/v1/ | Mirrored from PMXT, byte-identical | 2026-02-21T18 → 2026-04-16T05 |
pmxt/v2/ | Mirrored from PMXT, byte-identical | 2026-04-13T19 → 2026-08-09T23 |
third-party/ag6/ | One collector, no second machine to cross-check it. Mirrored 2026-08-26 from polymarket-archive.ag6.ai | 2026-08-01T13 → 2026-08-15T09 |
contrib/devpyle/ | Contributed; not witnessed by this archive. Recorded by devpyle, crypto Up-or-Down markets only, in his own schema. What that means | 2026-08-15T10 → 2026-08-18T05 |
v3/ | Our own capture, several machines, merged and checked hour by hour | 2026-08-18T06 → ongoing |
The four eras do not join up. AG6's contiguous run of 134 hours starts at 2026-08-09T20, four hours before PMXT's last hour rather than after it, so those two overlap; its listing begins earlier still, at 2026-08-01T13, with two hours that sit inside PMXT's span. After AG6 ends there is a gap of 68 hours, 2026-08-15T10 to 2026-08-18T05, that nothing here covers for the whole venue - the contributed tier holds only a handful of crypto markets for those hours - and then our own capture begins.