Polymarket Orderbook Archive archive.pendulumflow.com

Questions

New here? Start with what you can actually do with this data: ten uses, each with a worked example and the columns it needs.

Is it really free?

Yes, and there is no account and no key. Use it for anything, including commercially.

You do have to credit us. The archive is published under CC BY 4.0, and attribution is the one condition that licence carries: if you publish, sell or build on this data, say where it came from and link back to archive.pendulumflow.com. The licence requires it. Credit the collector of the era you used, which Credits sets out in full.

It is not free to produce. The archive was 2.8 TiB when counted on 2026-10-06, and a published v3 hour averages about 1.1 GB, measured over the same listing. Storage is billed every month whether anyone downloads a byte or not, and nobody here takes a wage from it. Donations go straight to that bill. Several machines record the feed around the clock, and running them is part of what the bill pays for. If you are building on this, tell your own users where it came from and that donations keep it running. A link is worth more to this archive than most single donations, and it is the only thing that reaches people we never will. There is a donation address on every page.

Is there a rate limit?

There wasn't. Then a few readers got greedy and started pulling so hard that the storage behind the archive hit its cap and turned everyone away with HTTP 429 errors, sometimes for minutes at a time. So we're adding a per-client limit at the edge. One heavy reader won't be able to knock the archive over for everyone else any more. If you do get a 429, keep the ranges you already have, wait 30 to 60 seconds, then carry on from where you stopped.

Do I have to download whole hours?

Under v3/, no. Every v3 hour ships a manifest.json that gives each event type its own byte range, so you can fetch just the trades from an hour with an HTTP range request and skip the rest. That range carries its own checksum, so a partial download is still verifiable. The v3 notes show the commands, and an AI assistant can write them for you from this paragraph.

The earlier eras have no manifest. pmxt/v1/, pmxt/v2/ and third-party/ag6/ are one Parquet file per hour and nothing else, so there is no byte range to ask for and an hour is the smallest unit. Parquet is columnar, so a reader that can seek over HTTP still fetches only the columns it needs, and you never download the whole file.

Why are two hours marked complete when one is half the size?

Because complete is a statement about time, not volume. It means all sixty minutes of that hour carry data and the first and last events fall within five seconds of the hour's edges. That is the check the audit runs on every published v3 hour. How much data an hour holds depends on how busy the market was, and a quiet Saturday hour can legitimately be a fraction of the size of a busy one.

What is missing?

A 68 hour hole between the AG6 era and ours, 2026-08-15T10 to 2026-08-18T05. AG6's archive stops at 2026-08-15T09 and our own capture begins at 2026-08-18T06. Nothing we have found covers the hours between. If you hold those hours we will publish them and credit you.

Four hours in the mirrored eras with no file at all, 2026-04-04T18, 2026-06-11T04, 2026-06-11T05, 2026-06-11T06, which is why the front page counts them among its missing hours. Two more hours exist as files but hold no rows, so the front page does not count them. The audit names all six. The two empty files are what PMXT published, mirrored as they are; for the four absent hours, PMXT's own origin was asked on 2026-09-02: it answered 404 for these hours and served every neighbouring hour we checked, so they are absent at source, not missed by our mirror.

The first days of our own capture are thinner than the rest. We were still hardening the pipeline. The earliest hours have no witness columns, the columns that record which machine heard each row, because those columns did not exist yet. Some later hours lack them too, so check the file you hold rather than assuming a cut-over date. Some of the early hours were heard by only one machine. That is real data with nothing to cross-check it against, and the audit marks it as such rather than averaging it into the rest. Later hours are heard by several machines, and each hour says which.

The audit is where all of this is counted, hour by hour, and it is where anything new lands: how many machines heard each hour, which hours are partial, and the capture losses we have recorded, with what we can and cannot say about each. We would rather tell you than have you find it.

How many different formats are there?

Three formats across four corpora. pmxt/v2/ and third-party/ag6/ are the same sixteen columns under different names, checked column by column, so a query written for one runs unchanged on the other. pmxt/v1/ is five columns with JSON inside, and v3/ is between 25 and 28 depending on the hour, 28 in the newest, measured over every hour published by 2026-10-06.

They differ for two reasons. Different collectors produced them, and rewriting the old ones to match the new one would mean publishing something other than what was recorded. And our own recording got more thorough: v3 adds a sequence that makes de-duplication exact, columns naming which machine heard each row and how late it arrived, and the market metadata the earlier eras never carried, so each hour can be checked.

The format notes cover each era column by column, and the working notes warn about the likeliest trap: v1 names things differently, so a query written for the later eras returns nothing on it and does not tell you.

Can I mirror the whole thing?

Yes. Start from a prefix's SHA256SUMS.txt, which lists every file under it, then verify what you got with the same file. The listing pages carry the commands.

If you host a copy, credit pendulumflow and link back to archive.pendulumflow.com, and carry the donation addresses this site publishes. Serving a mirror costs you bandwidth; capturing the feed that fills it costs us money every hour, and donations are the only thing paying that bill.

Will it keep running?

That depends partly on donations and partly on us, so we would rather tell you what does not depend on either. The licence means nobody needs our permission to keep a copy of anything already published. If this archive matters to your work, mirror it. Each prefix lists every file it holds, with checksums, precisely so a copy can be made and verified without asking us. We would call that a good outcome.

Something we have not answered, or got wrong? https://x.com/PendulumFlow

Answers to questions sent through the feedback form

If you sent us a question and didn't leave contact details, your answer will be here.

Range requests to dl.pendulumflow.com sometimes return HTTP 429 even with one connection and a minute between requests. Is this an origin-storage issue, and is there a mirror or bulk route for v3 manifest ranges?

Asked 2026-10-08

Thanks for such a precise report. You're right, and our FAQ was wrong about rate limits. We've fixed it.

It is an origin-storage limit, and nothing you're doing causes it. Our storage provider caps requests per account, and every reader shares that one cap. At busy times we see 300,000 to 560,000 range requests an hour across all readers. When that tips over the cap, the storage returns 429 to everyone for a few minutes. Over the last week that was between 0.4% and 2.6% of requests on any given day, almost all in short bursts. Your single connection with a minute's spacing was never the problem.

What you're doing is the right approach: keep completed ranges, back off 30 to 60 seconds on a 429 (no Retry-After is sent), and resume by range. Manifest ranges are the supported route. There isn't a second mirror yet.

We're working on two things. One is a per-client limit at our edge, so a single heavy reader can't push the whole account over the cap. The other is serving the files from somewhere without a shared per-account request cap. A bulk route for pulling the whole archive is on the list too. If that's what you need, tell us through the form (and leave a contact this time, so we can tell you when it's ready).

Credit the collector: pendulumflow for V3, PMXT for V1 and V2, AG6 for their V2 archive. Ours and PMXT's are CC BY 4.0; AG6 states no licence. How to cite.

Serving these bytes is not endorsing them. We are not affiliated with Polymarket.

For AI readers: llms.txt, what this archive holds and the questions it cannot answer.

https://x.com/PendulumFlow

Join the Pendulum Flow Discord: other people who build on prediction-market data, comparing notes and sharing tooling and findings.