Polymarket Orderbook Archive archive.pendulumflow.com

V3 data overview

DownloadAudit

v3/, live data from PendulumFlow, 2026-08-18T06 onward. An hour is a directory, v3/YYYY-MM-DD/HH/, holding one Parquet file and a manifest, and on recent hours a witness_stats.json timing file. Continues PMXT's v1 and v2; the collector is a rewrite of theirs.

How an hour is laid out

An hour is a single Parquet file. Its rows are grouped by event type, and the manifest gives each type its own byte range, row count and sha256.

That grouping is what makes a targeted question cheap. Parquet is columnar, so you never read the whole hour anyway; you read the columns you asked for. But because no row group straddles two event types, a reader who wants only trades can skip almost every row group, and can fetch just that byte range over HTTP without downloading the hour at all.

book

Full depth snapshots. bids and asks are nested lists of price and size. The only product carrying resting depth.

price_change

One row per changed price level. About 85% of an hour's bytes. Between two book snapshots these rows are the only record of depth changing.

best_bid_ask

Best bid, best ask and their difference, one row per change.

last_trade_price

Trades: price, size, side and the on-chain transaction_hash.

new_market

Market creation, question text, slug, outcomes, and the two complementary assets_ids.

market_resolved

Resolutions: winning asset and outcome.

tick_size_change

Tick size changes, old and new.

manifest.json

Provenance for the hour: row counts and a sha256 for the file and for every event type's byte range, the sequence range, and which capture hosts contributed. Small, and worth reading first. One field needs care. On some hours merger_git_sha holds an error string instead of a commit id. Check the shape of a field before using it, and treat every value as a string.

<hour>.witness_stats.json

How long each capture machine took to receive each kind of event: timestamp_received - timestamp, as a histogram with p50, p90 and p99, per machine and per event type. It counts every row that arrived, including rows from machines whose copy lost the merge and so is not in the Parquet at all - the Parquet keeps one copy of a row, and this counts every copy that arrived. Read it as capture timing, not as market activity: a machine being slower says nothing about the venue. Only recent hours have one. Earlier hours can gain one where the per-machine recordings still exist; for the 83 oldest hours they do not, see the note under Columns.

Columns

One schema across every event type: a column that does not apply to a row is null there rather than absent, so you can read the whole hour into one frame without unioning anything. The Present on label beside each column comes from the published SCHEMA.json.

event_typestring (dictionary)nullable

Which kind of event the row is. The hour's rows are grouped by this column and no row group straddles two of them, so a reader who wants only trades can skip almost every row group, or fetch just that byte range over HTTP without downloading the hour.

timestamp_receivedtimestamp[us, tz=UTC]nullable

When our capture host received the event. Microseconds, and this is the column to sort by: fine enough that events rarely tie, where a millisecond clock ties constantly.

sequenceuint64nullable

Unique within the hour, so it de-duplicates exactly. It is not a global order: it is collector-local provenance, and sequence values from different machines are not comparable.

timestamptimestamp[ms, tz=UTC]nullable

The exchange's own timestamp, in milliseconds. Coarser than our receive time by three orders of magnitude, and the same value across every machine that heard the event, which is what makes it usable as a sort key inside the merge.

marketbinarynullable

Condition id, as bytes. Cast to text on read.

asset_idbinarynullable

Outcome token id, as bytes.

best_biddecimal(9,4)nullable

The highest resting bid at that moment.

best_askdecimal(9,4)nullable

The lowest resting ask at that moment.

spreaddecimal(9,4)nullable

best_ask minus best_bid, computed by the venue.

bidslist of struct<price: decimal(9,4), size: decimal(18,6)>nullable

Resting bid depth, as a native nested list of price and size rather than a JSON string. You can read it straight into a dataframe with no parsing step, and a reader that pushes predicates down can filter inside it.

askslist of struct<price: decimal(9,4), size: decimal(18,6)>nullable

Resting ask depth, nested the same way as bids.

pricedecimal(9,4)nullable

The price of the level that changed, or of the fill.

sizedecimal(18,6)nullable

The size at that price after the change, or the size filled.

sidestringnullable

BUY or SELL, from the taker's point of view.

fee_rate_bpsuint16nullable

The fee charged on the fill, in basis points.

transaction_hashbinarynullable

The Polygon transaction the fill settled in.

idstringnullable

The venue's own id for the lifecycle event.

assets_idslist of binarynullable

The market's two complementary outcome tokens.

winning_asset_idbinarynullable

The outcome token that resolved true.

winning_outcomestringnullable

That token's label, as the venue wrote it.

outcomeslist of stringnullable

The outcome labels, in the same order as assets_ids.

questionstringnullable

The market question as the venue wrote it.

slugstringnullable

The market's URL slug on polymarket.com.

old_tick_sizedecimal(9,4)nullable

The tick size before the change.

new_tick_sizedecimal(9,4)nullable

The tick size after it.

witness_setstring (dictionary)nullable

Every machine that heard the row, written |a|b|c| with a bar at both ends so that matching a name as a substring cannot half-match a longer one.

arrival_skewint64nullable

The gap in microseconds between the earliest and latest arrival of the row across machines, and null where only one machine heard it: with one machine there is no gap to measure, which is not the same as a gap of zero.

source_witnessstring (dictionary)nullable

Which machine's copy of the row the merge kept. That is decided per hour: which copy the merge kept is a per-hour fact, read from the hour's merger stamp: fcbb2804 and later keep the earliest-received copy; earlier stamps keep the copy from the machine most recently added to the fleet. Under either rule it is a bad answer to "what share of the feed did each machine hear"; use witness_set for that.

28 columns. The types above were read from the footer of 2026-08-28T23, not copied from documentation, so they are what your own reader will report. nullable says whether the schema permits a null at all, which is not the same as how often one appears.

83 of the oldest hours will never carry the witness columns. They are worked out by comparing what each machine separately recorded, and for those hours the separate recordings are gone, so no later run can produce them. An absent value there is not a bug, not a machine that heard nothing, and not something still arriving: the hour itself is intact, its market data complete and checksummed like every other hour.

Reading across that boundary is fine unless you name a witness column, which needs union_by_name = true in DuckDB or a unified schema in pyarrow. Counted 2026-08-27 from the surviving per-machine recordings, which are not published here. Each affected hour says so in its own manifest, under witness_attribution.

v3.1 sidecars (optional)

Not published yet: no served hour carries these files. They have run on one capture machine as a canary (the witness-h canary, 24 hours from 2026-09-26T12 to 2026-09-27T11), and the publisher will begin including them behind a switch. Until it does, every hour reads exactly as described above. Full column lists appear here, read from SCHEMA.json, once served hours carry the files.

Two optional files per capture machine per hour, in a folder of their own inside the hour: v3/YYYY-MM-DD/HH/v31/<machine>/, where the machine is one of the letters in the hour's witnesses. They are an addition, not a new format: the hour's Parquet file, its products, their columns and the v3/ paths do not change, and code that ignores the folder reads a v3.1 hour exactly as it reads any other. They describe how each machine received the data, not the market.

v31/<machine>/frame_index.parquetoptional10 columns on the canary

How each row reached that machine. One row per event the machine captured in the hour (on the canary the count equalled the machine's own row count in every hour), keyed by that machine's sequence and event_type: which connection and session received it (conn_slot, session_no), the frame's position on that connection (frame_seq), where the row sat inside a frame that carried several events (parent_index, child_index, child_count), a monotonic receive clock that does not jump when the wall clock is corrected (recv_monotonic_ns), and the exchange's own book hash (exchange_hash; on the canary it was filled on book rows only). With it you can check the sequence integrity of what the machine received and find its drops. Size: about 70% of the size of the hour's products per machine (median 664 MB an hour on the canary).

v31/<machine>/connection_events.parquetoptional21 columns on the canary

The machine's own record of its connections: connects, disconnects and reconnects, dropped updates, clock corrections, and the gap intervals between losing a connection and recovering it, stated explicitly rather than left for you to infer from missing rows. One row per event: between 79 and 9,217 an hour on the canary. Size: tens of kilobytes an hour per machine (median 44 KB on the canary).

Both files are one machine's view. Their sequence is that machine's own numbering (see what sequence is not), so it lines up with the hour's rows only where that machine's copy is the one the merge kept. They are for checking our capture, not an ordering of the venue's events.

An hour whose manifest.json has no v31 key has no sidecars.

More than one machine is watching

Most archives are one collector writing files. This one runs several capture hosts in different data centres, all subscribed to the same exchange feed, and merges what they heard. That is the part worth explaining, because merging two recordings of the same stream is harder than it sounds and doing it badly is invisible afterwards.

Not every hour has the same number of witnesses. The number changes as machines join and retire, so it is published per hour rather than as a range: the audit lists it for every hour, and each hour's own report names the machines. An hour with a single witness is still real data; it simply has nothing to cross-check against, so treat it with a little more caution than the rest.

Why you cannot just concatenate

Every host hears the same event, so joining their files end to end would repeat almost every row. The obvious fix is to match rows by sequence number, and that does not work: each collector numbers its own stream, so the same exchange event carries different sequence numbers on different hosts. Sequence identifies a row within one feed. It says nothing across feeds.

What we match on instead

The key is the exchange's own content. Every field the exchange sent, hashed, with our two local additions left out: the time we received it, and our sequence number. Two hosts that heard the same event produce the same key.

Some events genuinely repeat. The exchange really does send the same content twice, and a merge that collapsed those would quietly delete real data. So each row also carries its occurrence number within its own feed, first copy, second copy and so on, and the match is on the pair. The second copy on one host pairs with the second copy on the other. The result is that a repeat seen by both hosts stays a repeat, while the same single event heard twice becomes one row.

The hour boundary, where latency shows up

The hosts do not receive at the same moment. The measured spread between them runs from about 55 to 874 milliseconds, which is nothing until an event lands within a second of the top of the hour. Then one host files it in this hour and the other files it in the last one, and a naive merge writes it twice, once in each.

So each merged hour leaves behind the keys from its own tail, and the next hour checks against them before writing. It is occurrence-aware for the same reason as the join: if the event really happened twice, the second one survives. The count of rows dropped this way is recorded per hour rather than silently applied.

What is new in v3

What sequence is not

It is not a guaranteed global ordering of the hour. sequence is collector-local provenance: each capture host numbers what it sees, and where several hosts saw the same event the merged row carries one host's value, chosen by that hour's kept-copy rule (above). Numbers from different hosts are therefore not strictly comparable.

In practice it comes close. Sorting one hour's trades by sequence put timestamp_received in order with 10 inversions across 60,967 adjacent pairs, about 0.02%. But that is a description of how the collectors happened to run, not a property you can rely on. Use it to de-duplicate, not to order. What to order by is below.

The order in the file is not the order anyone saw

Rows are sorted rather than left as they arrived. Event type first, so each one is a contiguous byte range you can fetch on its own, and then within each event type by market, asset_id, the venue's own timestamp, and sequence. Every hour's manifest records that under layout and per-product order_by, so you can read it off the file rather than trusting this page. It is deterministic: build the same hour twice on two different machines and you get the same bytes.

A live subscriber watches one socket, with every event type interleaved across every market, in whatever order their connection happened to deliver. This file gives you the opposite: all of one event type together, one market at a time. Nobody ever saw the order these files are in, including us.

There is also no single arrival order for us to preserve. Several hosts hear the same feed and do not agree about ordering; the spread between them ran from 55 to 874 milliseconds when measured (see the hour boundary above). Keeping one host's order would mean discarding the rows only the other hosts heard, which is the entire reason for merging in the first place.

What that means for a backtest. If your strategy depends on the order events reached one particular machine, this archive cannot give you that, and no merged archive can. Sort by timestamp_received to approximate arrival, or by timestamp for the venue's own view of when things happened. Neither is a replay of a single socket, and you should not treat either as one. Nor does any row say what an order of yours would have received: a replay that fills the moment a price is touched overstates fills and understates their cost, so treat a touch-fill result as an upper bound, see A fill, from a price you touched.

Per-row provenance. 954 of the 1179 published hours carry it, in source_witness - but not contiguously, so there is no hour from which you can assume it is present. Check the file you hold.

What the manifest has always recorded, and still does, is which hosts contributed to the hour and how many rows each supplied, broken down by event type. That gives you the shape of the merge whether or not the per-row column is present in the file you are holding.

An hour is a container, not a claim about when events happened

Files are partitioned by timestamp_received. When our collector received the event, not when the venue reported it. Those usually differ by milliseconds. In principle they can differ by much more: if a capture host stalls and then recovers, its backlog drains into whichever hour is being written at the time, so an hour's file could carry events that occurred well before it. If you need event time, use timestamp and not the file's name.

We have not seen it happen. The one hour where we expected it, written just after a host recovered from a two-and-a-half-hour stall, was checked against event time and came back 100% in-hour on both event types that could have shown otherwise. The mechanism is real by design; we have not observed it.

How to check what you downloaded

Two recipes. Both use standard tools and neither needs anything from us but the manifest.

1. The whole file: the one you should normally use

sha256sum <date>T<hour>.parquet
# compare with .sha256 in that hour's manifest.json

One command and no Parquet awareness. The same command works on pmxt/v1/ and pmxt/v2/, where an hour is also a single file; there compare against SHA256SUMS.txt, because those eras have no manifest. This is the primary check.

2. One product, without downloading the hour

The manifest gives every product a byte range and its own digest, so you can verify a part of the file after fetching only that part:

# take byte_range and sha256 for the product from that hour's manifest.json
# the range is [start, end); HTTP Range is inclusive, so ask for end - 1
curl -r <start>-<end-1> https://dl.pendulumflow.com/v3/<date>/<hour>/<file>.parquet | sha256sum
# compare with .products.<product>.sha256

Note the end byte is one less than the manifest's. The manifest records a half-open range and HTTP Range is inclusive.

What the range check does and does not establish

A range digest checks bytes. It proves the bytes at those offsets are the bytes we recorded, and nothing else.

That those particular bytes are the last_trade_price rows is a claim made by the manifest. You do not have to take it on trust: the Parquet footer lists every row group's offsets and its event_type statistics, so you can read the footer and confirm the range covers exactly the row groups whose event type matches, and the footer is inside the file, so the whole-file digest already covers it. A wrong boundary in the manifest cannot make altered bytes pass a check; it can only mislabel which checked bytes you were given, and reading the footer detects that.

If you want no Parquet awareness at all, use recipe 1.

What a checksum here does and does not prove

Every published sha256 lets you check that the file you downloaded is the file we recorded. That catches a truncated download, a corrupted object, a proxy that rewrote something. It is worth doing and the digests exist so that you can.

It does not prove that neither the file nor its manifest was replaced by someone able to change both. A digest published beside the data it describes verifies consistency, not history. We do not currently sign manifests or anchor their digests anywhere outside this archive. Read these checksums as a corruption and consistency check. They are not evidence that the archive has never been altered.

Corrections

Two things earlier versions of this page said, and withdrew. That sequence was a guaranteed global ordering of the hour: it is collector-local, as the section on it now says. And that we had measured hour-boundary contamination after a host recovered from a stall: that was an inference, the direct measurement came back 100% in-hour, and the claim is withdrawn.

Pendulum Flow Discord

Something wrong with an hour, or a file you need that is not here? Say so in the Discord, where the people who run the site and prediction market data geeks compare notes.

Join the Discord

Credit the collector: pendulumflow for V3, PMXT for V1 and V2, AG6 for their V2 archive. Ours and PMXT's are CC BY 4.0; AG6 states no licence. How to cite.

Serving these bytes is not endorsing them. We are not affiliated with Polymarket.

For AI readers: llms.txt, what this archive holds and the questions it cannot answer.

https://x.com/PendulumFlow

Join the Pendulum Flow Discord: other people who build on prediction-market data, comparing notes and sharing tooling and findings.