V3 data overview
v3/, live data from PendulumFlow, 2026-08-18T06 onward. An hour is a
directory, v3/YYYY-MM-DD/HH/, holding one Parquet file and a manifest, and on recent hours a witness_stats.json timing file.
Continues PMXT's v1 and v2; the collector is a rewrite of theirs.
How an hour is laid out
An hour is a single Parquet file. Its rows are grouped by event type, and the manifest gives each type its own byte range, row count and sha256.
That grouping is what makes a targeted question cheap. Parquet is columnar, so you never read the whole hour anyway; you read the columns you asked for. But because no row group straddles two event types, a reader who wants only trades can skip almost every row group, and can fetch just that byte range over HTTP without downloading the hour at all.
book
Full depth snapshots. bids and asks are nested lists of price and size. The only product carrying resting depth.
price_change
One row per changed price level. About 85% of an hour's bytes. Between two book snapshots these rows are the only record of depth changing.
best_bid_ask
Best bid, best ask and their difference, one row per change.
last_trade_price
Trades: price, size, side and the on-chain transaction_hash.
new_market
Market creation, question text, slug, outcomes, and the two complementary assets_ids.
market_resolved
Resolutions: winning asset and outcome.
tick_size_change
Tick size changes, old and new.
manifest.json
Provenance for the hour: row counts and a sha256 for the file and for every event type's byte range, the sequence range, and which capture hosts contributed. Small, and worth reading first. One field needs care. On some hours merger_git_sha holds an error string instead of a commit id. Check the shape of a field before using it, and treat every value as a string.
<hour>.witness_stats.json
How long each capture machine
took to receive each kind of event: timestamp_received - timestamp, as a
histogram with p50, p90 and p99, per machine and per event type. It counts
every row that arrived, including rows from machines whose copy lost the
merge and so is not in the Parquet at all - the Parquet keeps one copy of a row, and
this counts every copy that arrived. Read it as capture timing, not as market activity: a
machine being slower says nothing about the venue. Only recent hours have
one. Earlier hours can gain one where the per-machine recordings still exist; for
the 83 oldest hours they do not, see the note under
Columns.
Columns
One schema across every event type: a column that does not apply to a row is null there
rather than absent, so you can read the whole hour into one frame without unioning
anything. The Present on label beside each column comes from the published
SCHEMA.json.
event_typestring (dictionary)nullable
Which kind of event the row is. The hour's rows are grouped by this column and no row group straddles two of them, so a reader who wants only trades can skip almost every row group, or fetch just that byte range over HTTP without downloading the hour.
timestamp_receivedtimestamp[us, tz=UTC]nullable
When our capture host received the event. Microseconds, and this is the column to sort by: fine enough that events rarely tie, where a millisecond clock ties constantly.
sequenceuint64nullable
Unique within the hour, so it de-duplicates exactly. It is not a global order: it is collector-local provenance, and sequence values from different machines are not comparable.
timestamptimestamp[ms, tz=UTC]nullable
The exchange's own timestamp, in milliseconds. Coarser than our receive time by three orders of magnitude, and the same value across every machine that heard the event, which is what makes it usable as a sort key inside the merge.
marketbinarynullable
Condition id, as bytes. Cast to text on read.
asset_idbinarynullable
Outcome token id, as bytes.
best_biddecimal(9,4)nullable
The highest resting bid at that moment.
best_askdecimal(9,4)nullable
The lowest resting ask at that moment.
spreaddecimal(9,4)nullable
best_ask minus best_bid, computed by the venue.
bidslist of struct<price: decimal(9,4), size: decimal(18,6)>nullable
Resting bid depth, as a native nested list of price and size rather than a JSON string. You can read it straight into a dataframe with no parsing step, and a reader that pushes predicates down can filter inside it.
askslist of struct<price: decimal(9,4), size: decimal(18,6)>nullable
Resting ask depth, nested the same way as bids.
pricedecimal(9,4)nullable
The price of the level that changed, or of the fill.
sizedecimal(18,6)nullable
The size at that price after the change, or the size filled.
sidestringnullable
BUY or SELL, from the taker's point of view.
fee_rate_bpsuint16nullable
The fee charged on the fill, in basis points.
transaction_hashbinarynullable
The Polygon transaction the fill settled in.
idstringnullable
The venue's own id for the lifecycle event.
assets_idslist of binarynullable
The market's two complementary outcome tokens.
winning_asset_idbinarynullable
The outcome token that resolved true.
winning_outcomestringnullable
That token's label, as the venue wrote it.
outcomeslist of stringnullable
The outcome labels, in the same order as assets_ids.
questionstringnullable
The market question as the venue wrote it.
slugstringnullable
The market's URL slug on polymarket.com.
old_tick_sizedecimal(9,4)nullable
The tick size before the change.
new_tick_sizedecimal(9,4)nullable
The tick size after it.
witness_setstring (dictionary)nullable
Every machine that heard the row, written |a|b|c| with a bar at both
ends so that matching a name as a substring cannot half-match a longer one.
arrival_skewint64nullable
The gap in microseconds between the earliest and latest arrival of the row across machines, and null where only one machine heard it: with one machine there is no gap to measure, which is not the same as a gap of zero.
source_witnessstring (dictionary)nullable
Which machine's copy of the row the merge kept. That is decided per hour:
which copy the merge kept is a per-hour fact, read from the hour's merger stamp: fcbb2804 and later keep the earliest-received copy; earlier stamps keep the copy from the machine most recently added to the fleet. Under either rule it is a bad answer to "what share of the feed did
each machine hear"; use witness_set for that.
28 columns. The types above were read from
the footer of 2026-08-28T23,
not copied from documentation, so they are what your own reader will report.
nullable says whether the schema permits a null at all, which is not the same as
how often one appears.
83 of the oldest hours will never carry the witness columns. They are worked out by comparing what each machine separately recorded, and for those hours the separate recordings are gone, so no later run can produce them. An absent value there is not a bug, not a machine that heard nothing, and not something still arriving: the hour itself is intact, its market data complete and checksummed like every other hour.
Reading across that boundary is fine unless you name a witness column,
which needs union_by_name = true in DuckDB or a unified schema in pyarrow.
Counted 2026-08-27 from the surviving per-machine recordings, which are not
published here. Each affected hour says so in its own manifest, under
witness_attribution.
v3.1 sidecars (optional)
Not published yet: no served hour carries these files.
They have run on one capture machine as a canary (the witness-h canary, 24 hours from 2026-09-26T12 to 2026-09-27T11), and the publisher
will begin including them behind a switch. Until it does, every hour reads exactly as
described above. Full column lists appear here, read from
SCHEMA.json, once served hours carry the files.
Two optional files per capture machine per hour, in a folder of their
own inside the hour: v3/YYYY-MM-DD/HH/v31/<machine>/, where
the machine is one of the letters in the hour's witnesses. They are an
addition, not a new format: the hour's Parquet file, its products, their columns and the
v3/ paths do not change, and code that ignores the folder reads a v3.1 hour
exactly as it reads any other. They describe how each machine received the data, not the
market.
v31/<machine>/frame_index.parquetoptional10 columns on the canary
How each row reached that machine. One row per event the machine captured in
the hour (on the canary the count equalled the machine's own row count in every hour), keyed
by that machine's sequence and event_type: which connection and
session received it (conn_slot, session_no), the frame's position
on that connection (frame_seq), where the row sat inside a frame that carried
several events (parent_index, child_index,
child_count), a monotonic receive clock that does not jump when the wall clock
is corrected (recv_monotonic_ns), and the exchange's own book hash
(exchange_hash; on the canary it was filled on book rows only).
With it you can check the sequence integrity of what the machine received and find its
drops. Size: about 70% of the size of the hour's products per machine (median 664 MB an hour on the canary).
v31/<machine>/connection_events.parquetoptional21 columns on the canary
The machine's own record of its connections: connects, disconnects and reconnects, dropped updates, clock corrections, and the gap intervals between losing a connection and recovering it, stated explicitly rather than left for you to infer from missing rows. One row per event: between 79 and 9,217 an hour on the canary. Size: tens of kilobytes an hour per machine (median 44 KB on the canary).
Both files are one machine's view. Their sequence is that machine's own
numbering (see what sequence is not), so it lines up with the
hour's rows only where that machine's copy is the one the merge kept. They are for checking
our capture, not an ordering of the venue's events.
An hour whose manifest.json has no v31
key has no sidecars.
More than one machine is watching
Most archives are one collector writing files. This one runs several capture hosts in different data centres, all subscribed to the same exchange feed, and merges what they heard. That is the part worth explaining, because merging two recordings of the same stream is harder than it sounds and doing it badly is invisible afterwards.
Not every hour has the same number of witnesses. The number changes as machines join and retire, so it is published per hour rather than as a range: the audit lists it for every hour, and each hour's own report names the machines. An hour with a single witness is still real data; it simply has nothing to cross-check against, so treat it with a little more caution than the rest.
Why you cannot just concatenate
Every host hears the same event, so joining their files end to end would repeat almost every row. The obvious fix is to match rows by sequence number, and that does not work: each collector numbers its own stream, so the same exchange event carries different sequence numbers on different hosts. Sequence identifies a row within one feed. It says nothing across feeds.
What we match on instead
The key is the exchange's own content. Every field the exchange sent, hashed, with our two local additions left out: the time we received it, and our sequence number. Two hosts that heard the same event produce the same key.
Some events genuinely repeat. The exchange really does send the same content twice, and a merge that collapsed those would quietly delete real data. So each row also carries its occurrence number within its own feed, first copy, second copy and so on, and the match is on the pair. The second copy on one host pairs with the second copy on the other. The result is that a repeat seen by both hosts stays a repeat, while the same single event heard twice becomes one row.
The hour boundary, where latency shows up
The hosts do not receive at the same moment. The measured spread between them runs from about 55 to 874 milliseconds, which is nothing until an event lands within a second of the top of the hour. Then one host files it in this hour and the other files it in the last one, and a naive merge writes it twice, once in each.
So each merged hour leaves behind the keys from its own tail, and the next hour checks against them before writing. It is occurrence-aware for the same reason as the join: if the event really happened twice, the second one survives. The count of rows dropped this way is recorded per hour rather than silently applied.
What is new in v3
- A sound de-duplication key. Every row carries a
sequence, and on 2026-08-24T05 it was unique across the hour: 84,148,495 rows, zero duplicates. De-duplicating by sequence needs no tie-breaking on timestamps, which v1 and v2 do need. - Sub-millisecond receive times.
timestamp_receivedis microseconds here, against milliseconds in v1 and v2. The exchange's owntimestampstays milliseconds in every era, so the finer resolution is ours, not Polymarket's, and it measures when we received an event, not when it happened. - Checked coverage. A separate pass counts each hour's events and labels the hour complete or partial; every partial hour is named on the audit.
- Checksums in each hour's manifest, for the whole file and for each event type's byte range, and in SHA256SUMS.txt.
What sequence is not
It is not a guaranteed global ordering of the hour. sequence
is collector-local provenance:
each capture host numbers what it sees, and where several hosts saw the same event the
merged row carries one host's value, chosen by that hour's kept-copy rule (above). Numbers
from different hosts are therefore not strictly comparable.
In practice it comes close. Sorting one hour's trades by sequence put
timestamp_received in order with 10 inversions across 60,967 adjacent pairs,
about 0.02%. But that is a description of how the collectors happened to run, not a
property you can rely on. Use it to de-duplicate, not to order. What to
order by is below.
The order in the file is not the order anyone saw
Rows are sorted rather than left as they arrived. Event type first, so
each one is a contiguous byte range you can fetch on its own, and then within each event
type by market, asset_id, the venue's own timestamp, and
sequence. Every hour's manifest records that under layout and
per-product order_by, so you can read it off the file rather than trusting
this page. It is deterministic: build the same hour twice on two different machines and
you get the same bytes.
A live subscriber watches one socket, with every event type interleaved across every market, in whatever order their connection happened to deliver. This file gives you the opposite: all of one event type together, one market at a time. Nobody ever saw the order these files are in, including us.
There is also no single arrival order for us to preserve. Several hosts hear the same feed and do not agree about ordering; the spread between them ran from 55 to 874 milliseconds when measured (see the hour boundary above). Keeping one host's order would mean discarding the rows only the other hosts heard, which is the entire reason for merging in the first place.
What that means for a backtest. If your strategy depends on the order
events reached one particular machine, this archive cannot give you that, and no merged
archive can. Sort by timestamp_received to approximate arrival, or by
timestamp for the venue's own view of when things happened. Neither is a
replay of a single socket, and you should not treat either as one. Nor does any row say
what an order of yours would have received: a replay that fills the moment a price is
touched overstates fills and understates their cost, so treat a touch-fill result as an
upper bound, see A fill, from a price you touched.
Per-row provenance. 954 of the 1179 published hours carry it, in
source_witness - but not contiguously, so there is no
hour from which you can assume it is present. Check the file you hold.
What the manifest has always recorded, and still does, is which hosts contributed to the hour and how many rows each supplied, broken down by event type. That gives you the shape of the merge whether or not the per-row column is present in the file you are holding.
An hour is a container, not a claim about when events happened
Files are partitioned by timestamp_received. When our collector received
the event, not when the venue reported it. Those usually differ by milliseconds. In
principle they can differ by much more: if a capture host stalls and then recovers, its
backlog drains into whichever hour is being written at the time, so an hour's file could
carry events that occurred well before it. If you need event time, use
timestamp and not the file's name.
We have not seen it happen. The one hour where we expected it, written just after a host recovered from a two-and-a-half-hour stall, was checked against event time and came back 100% in-hour on both event types that could have shown otherwise. The mechanism is real by design; we have not observed it.
How to check what you downloaded
Two recipes. Both use standard tools and neither needs anything from us but the manifest.
1. The whole file: the one you should normally use
sha256sum <date>T<hour>.parquet
# compare with .sha256 in that hour's manifest.json
One command and no Parquet awareness. The same command works on
pmxt/v1/ and pmxt/v2/, where an hour is also a single file; there
compare against SHA256SUMS.txt, because those eras have no manifest. This is
the primary check.
2. One product, without downloading the hour
The manifest gives every product a byte range and its own digest, so you can verify a part of the file after fetching only that part:
# take byte_range and sha256 for the product from that hour's manifest.json
# the range is [start, end); HTTP Range is inclusive, so ask for end - 1
curl -r <start>-<end-1> https://dl.pendulumflow.com/v3/<date>/<hour>/<file>.parquet | sha256sum
# compare with .products.<product>.sha256
Note the end byte is one less than the manifest's. The manifest
records a half-open range and HTTP Range is inclusive.
What the range check does and does not establish
A range digest checks bytes. It proves the bytes at those offsets are the bytes we recorded, and nothing else.
That those particular bytes are the last_trade_price rows is a
claim made by the manifest. You do not have to take it on trust: the Parquet footer lists
every row group's offsets and its event_type statistics, so you can read the
footer and confirm the range covers exactly the row groups whose event type matches, and
the footer is inside the file, so the whole-file digest already covers it. A wrong
boundary in the manifest cannot make altered bytes pass a check; it can only mislabel
which checked bytes you were given, and reading the footer detects that.
If you want no Parquet awareness at all, use recipe 1.
What a checksum here does and does not prove
Every published sha256 lets you check that the file you downloaded is the file we recorded. That catches a truncated download, a corrupted object, a proxy that rewrote something. It is worth doing and the digests exist so that you can.
It does not prove that neither the file nor its manifest was replaced by someone able to change both. A digest published beside the data it describes verifies consistency, not history. We do not currently sign manifests or anchor their digests anywhere outside this archive. Read these checksums as a corruption and consistency check. They are not evidence that the archive has never been altered.
Corrections
Two things earlier versions of this page said, and withdrew. That
sequence was a guaranteed global ordering of the hour: it is collector-local,
as the section on it now says. And that we had measured hour-boundary contamination after a
host recovered from a stall: that was an inference, the direct measurement came back 100%
in-hour, and the claim is withdrawn.
Pendulum Flow Discord
Something wrong with an hour, or a file you need that is not here? Say so in the Discord, where the people who run the site and prediction market data geeks compare notes.
Join the Discord