Assymetrix Launches the Deepest Independent Prediction Market Data APIs
Read more
Read more
Assymetrix Launches the Deepest Independent Prediction Market Data APIs
Read more
Read more
Skip the Decoder: Polymarket Historical Data for Developers and Quants
Skip the Decoder: Polymarket Historical Data for Developers and Quants
Skip the Decoder: Polymarket Historical Data for Developers and Quants
Polymarket historical data for developers and quants: fetch, decode, normalize pipeline, trade/wallet/candle schemas, and a pre-normalized API backed by...

Skip the Decoder: Polymarket Historical Data for Developers and Quants
The fastest path to usable Polymarket history is a normalized data API for teams that need scale without engineering overhead, or the community parquet dataset paired with a custom decoding pipeline for teams needing full control over raw fields. Start with the SII-WANGZJ/Polymarket_data repository for bulk files and CLI tools, or a normalized feed like data.assymetrix.com/polymarket if you want clean tables on day one.
TL;DR:
Using the normalized data API provides ready-to-use tables, saving weeks of on-chain decoding work and eliminating manual filtering of noise from contracts like NegRisk.
Community datasets like Polymarket_data offer bulk parquet files that help identify schema issues early and reduce initial setup time but require ongoing maintenance of custom decode pipelines.
Raw on-chain logs require resolving condition IDs, filtering noise from wrapper contracts, and normalizing token perspectives, making direct fetch strategies complex and error-prone.
The most useful datasets for research include trade fills, wallet activity, and OHLCV candles, with a recommended snapshot cadence matching the shortest timeframe needed for your analysis.
Assymetrix’s API covers over 900 million events and simplifies deep Polymarket data access, allowing quick integration without building or maintaining a decoder yourself.
Assymetrixassymetrix.comUse Polymarket Data Without DecodingAssymetrix provides normalized Polymarket data through one API, helping developers and quants access structured historical market data at scale.Explore Assymetrix
Table of Contents
Building the pipeline: fetch, decode, normalize, store, query
Choosing your data source: on-chain, datasets, or APIs
Why raw on-chain decoding breaks pipelines
Schemas worth building: trades, wallets, markets, candles
What deep history actually unlocks
Keeping a pipeline trustworthy over time
How Assymetrix removes the decoding burden
Build versus buy: a practical decision checklist
Getting started with the Assymetrix Data API
Where to go next for datasets and docs
Sources
Building the pipeline: fetch, decode, normalize, store, query
Every serious Polymarket data project follows the same five stages, whether built in house or bought as a service. Fetching pulls raw events from a Polygon RPC endpoint, the Gamma API, or a bulk CLI export. Decoding turns logs into structured records: order fills, market resolutions, position transfers. Normalization applies consistent units, timestamps, and a single outcome perspective across markets. Storage lands the result in parquet files or a time-series database sized for the query patterns you expect. Querying is the layer your bots, backtests, and dashboards actually touch.
A minimal build order looks like this:
Stand up an indexer against Polygon RPC or a bulk export CLI and checkpoint block heights.
Write decode logic for order-filled and resolution events, with retries on RPC timeouts.
Normalize outcome tokens, wallet addresses, and price units into a single schema.
Write to parquet or a columnar store, partitioned by market and date.
Expose a query layer for backtests, dashboards, or live bot signals.
Snapshot cadence matters more than most teams expect. Sampling the CLOB orderbook every few seconds versus every minute changes what microstructure questions you can answer later, so decide cadence before you backfill, not after.
Choosing your data source: on-chain, datasets, or APIs
Three practical routes exist for getting Polymarket history into a usable form, and each fits a different resource profile.
Direct Polygon RPC fetch: full control over raw events, but you own log parsing, reorg handling, and rate limit management indefinitely.
Community bulk datasets: the SII-WANGZJ/Polymarket_data project ships prebuilt orderfilled, trades, markets, quant, and users parquet files with CLI commands for fetching and decoding on-chain blocks, which saves the initial backfill work.
CLOB V2 SDKs: built primarily for order placement and execution, they expose live orderbook state but are not designed as a historical archive, so pairing them with a separate history source is standard.
Normalized data APIs: trade the flexibility of raw access for pre-cleaned tables, cutting integration time from weeks to hours.
A technical breakdown of Polymarket’s API layers notes that the platform’s data API covers user-level analytics, including full trade history with timestamps and sizes, settled on Polygon, which is useful context when deciding how much of the raw chain you actually need to touch yourself.
Pro Tip: Pull a small time slice from a community dataset before committing to a full backfill: it exposes schema quirks and gaps faster than reading documentation alone.

Why raw on-chain decoding breaks pipelines
Polymarket settles trades on Polygon, and the on-chain event log alone does not tell you which human decision produced a given fill. Several decoding steps trip up pipelines that look correct on the surface.
Condition ID resolution: each market’s condition ID must be mapped to human-readable metadata such as question text and resolution date, and that mapping lives outside the raw trade logs.
NegRisk contract noise: multi-outcome markets route trades through NegRisk wrapper contracts, and filtering out contract-level noise from genuine user trades takes deliberate rules, not a simple address check.
Token perspective: every market has YES and NO outcome tokens, and mixing perspectives without normalizing to one side produces inconsistent price and volume series across markets.
Maker and taker signs: fills need consistent signed volume conventions, or aggregated P&L and flow metrics silently invert.
Reorg protection: blocks can be reorganized shortly after being produced, so writes need a confirmation buffer before being treated as final.
None of these are exotic problems, but skipping any one of them produces a dataset that looks complete and is quietly wrong.
Schemas worth building: trades, wallets, markets, candles
Four tables cover most backtesting and research needs and getting their fields right up front saves rework later.
Table | Key fields | Purpose |
|---|---|---|
trades | timestamp, block, market_id, outcome_id, wallet, side, price, size, maker_taker, tx_hash | Trade-level fills for replay and P&L |
users | wallet, split_type, position_side, cumulative_volume | Wallet-level activity and position tracking |
quant | market_id, timestamp, yes_price, yes_size, signed_volume | Unified YES-perspective series for cross-market modeling |
ohlcv | market_id, interval_start, open, high, low, close, volume | Model inputs at fixed time intervals |
The SII-WANGZJ/Polymarket_data project’s quant.parquet and users.parquet files follow this general pattern, splitting wallet activity by position type and unifying prices to a single outcome perspective so you are not reconciling YES and NO series by hand. Pick a snapshot cadence for OHLCV candles that matches your shortest research question: a one-minute cadence supports microstructure work that a five-minute cadence cannot.
What deep history actually unlocks
Once trades, wallets, and candles are normalized, four workflows become straightforward.
Backtesting: replay trade-level fills in timestamp order, aligning each fill against the OHLCV snapshot in force at that moment to simulate realistic entry and exit prices.
Smart Money scoring: compute a rolling P&L per wallet, then weight it by trade frequency and position persistence, since a wallet with consistent, sized positions over time is a stronger signal than one large lucky bet.
ML feature generation: derive time-window aggregates, rolling volume, price volatility, and orderbook imbalance proxies as inputs for classification or ranking models.
Microstructure research: study spread behavior, order flow imbalance, and price impact around resolution events using dense CLOB snapshots.
Each of these depends on the same underlying normalization work described above. A guide to backtesting prediction market strategies walks through replay logic in more depth, and separate coverage of identifying Smart Money wallets on Polymarket details wallet scoring methods beyond the simple version above.
Keeping a pipeline trustworthy over time
A pipeline that works on day one can quietly drift once markets change shape or volume grows.
Write unit tests for decode logic against known historical blocks with expected output.
Require multiple block confirmations before treating a write as final, and add reorg detection on top.
Monitor daily event counts against a rolling baseline to catch silent gaps early.
Test new decode logic against replay data before pointing it at a live feed.
Keep a hot store for recent data your bots query constantly and a cold store for full history you query rarely.
Pro Tip: Log the raw event count alongside the decoded record count for every batch: a growing gap between the two numbers is usually the first sign of a decoding bug.
How Assymetrix removes the decoding burden
Assymetrix indexes Polymarket on-chain activity going back to September 2020, covering more than 900 million events and over 200 million OHLCV snapshots, all delivered as normalized tables rather than raw logs.
Polygon RPC management, condition ID mapping, and NegRisk filtering are handled externally.
Trade fills, wallet activity, and resolution events are delivered pre-normalized to a single outcome perspective.
API calls typically request a market’s trade history or candle series over a date range and return clean rows ready for a dataframe.
Documentation for Polymarket-specific endpoints lives at data.assymetrix.com/polymarket, alongside broader Python integration examples.
This does not replace the value of understanding the underlying decode logic covered above. It replaces the need to maintain it yourself in production.
Build versus buy: a practical decision checklist
The right answer depends on four questions, not on preference. How novel is your research question: does it need raw, unfiltered fields, or standard OHLCV and trade tables? How much fidelity do you need on token perspective and maker/taker signs? Does your team have bandwidth to maintain a decoder through Polymarket’s contract changes? And how fast do you need results, this week or this quarter? Institutional research also needs reproducibility: a documented, versioned dataset is easier to audit than a hand-rolled pipeline nobody remembers the details of a year later.
— Dean
Getting started with the Assymetrix Data API
Raw Polymarket history is scattered across on-chain logs, wrapper contracts, and shifting condition IDs, and building a reliable decoder is a real engineering project, not a weekend script. The Assymetrix Data API at data.assymetrix.com/api skips that work entirely, delivering normalized trade fills, wallet activity, resolution data, and OHLCV candles across more than 900 million indexed events dating back to September 2020.

Request an API key to pull historical trade and candle data directly into your environment.
Try the sandbox against a single market before committing to a full historical backfill.
Request a bulk export if your research needs the full history in one file rather than paginated calls.
Full documentation and endpoint references for Polymarket coverage live at data.assymetrix.com/polymarket, with the main product and sign-up flow at the Data API page.
Where to go next for datasets and docs
The SII-WANGZJ/Polymarket_data repository for bulk parquet files and CLI decoding tools.
The Chainstack developer guide for API layer and CLOB V2 details.
Assymetrix’s Polymarket analytics guide for the full on-chain to analysis-ready workflow.
Sources
FAQ
What does Polymarket historical data actually include?
It spans trade-level fills, CLOB orderbook snapshots, wallet-level activity, market resolution events, and OHLCV candles built from those fills. Most research projects need at least trades and candles, with wallet data added for Smart Money analysis.
Why is fetching raw Polymarket data directly so complex?
Raw data lives as on-chain logs on Polygon, and turning those into usable rows means resolving condition IDs to market metadata, filtering NegRisk contract noise from genuine trades, and normalizing YES and NO token perspectives consistently. Skipping any of these steps tends to produce data that looks right but is not.
Should I use the community Polymarket_data dataset or an API?
The SII-WANGZJ/Polymarket_data dataset is a strong starting point for teams that want raw parquet files and are comfortable maintaining their own decode logic going forward. A normalized API is the faster route when you want clean tables without owning that maintenance.
What can I build with deep Polymarket history?
Trade-level and candle history support strategy backtesting, Smart Money wallet scoring, machine learning feature generation, and market microstructure research. Each of these depends on having consistent, normalized fields rather than raw, unfiltered logs.
Does Assymetrix require blockchain decoding on my end?
No. The Assymetrix Data API delivers pre-normalized Polymarket trade fills, wallet activity, resolution data, and OHLCV candles across more than 900 million indexed on-chain events dating back to September 2020, so your team queries clean tables instead of parsing logs.
Skip the Decoder: Polymarket Historical Data for Developers and Quants
The fastest path to usable Polymarket history is a normalized data API for teams that need scale without engineering overhead, or the community parquet dataset paired with a custom decoding pipeline for teams needing full control over raw fields. Start with the SII-WANGZJ/Polymarket_data repository for bulk files and CLI tools, or a normalized feed like data.assymetrix.com/polymarket if you want clean tables on day one.
TL;DR:
Using the normalized data API provides ready-to-use tables, saving weeks of on-chain decoding work and eliminating manual filtering of noise from contracts like NegRisk.
Community datasets like Polymarket_data offer bulk parquet files that help identify schema issues early and reduce initial setup time but require ongoing maintenance of custom decode pipelines.
Raw on-chain logs require resolving condition IDs, filtering noise from wrapper contracts, and normalizing token perspectives, making direct fetch strategies complex and error-prone.
The most useful datasets for research include trade fills, wallet activity, and OHLCV candles, with a recommended snapshot cadence matching the shortest timeframe needed for your analysis.
Assymetrix’s API covers over 900 million events and simplifies deep Polymarket data access, allowing quick integration without building or maintaining a decoder yourself.
Assymetrixassymetrix.comUse Polymarket Data Without DecodingAssymetrix provides normalized Polymarket data through one API, helping developers and quants access structured historical market data at scale.Explore Assymetrix
Table of Contents
Building the pipeline: fetch, decode, normalize, store, query
Choosing your data source: on-chain, datasets, or APIs
Why raw on-chain decoding breaks pipelines
Schemas worth building: trades, wallets, markets, candles
What deep history actually unlocks
Keeping a pipeline trustworthy over time
How Assymetrix removes the decoding burden
Build versus buy: a practical decision checklist
Getting started with the Assymetrix Data API
Where to go next for datasets and docs
Sources
Building the pipeline: fetch, decode, normalize, store, query
Every serious Polymarket data project follows the same five stages, whether built in house or bought as a service. Fetching pulls raw events from a Polygon RPC endpoint, the Gamma API, or a bulk CLI export. Decoding turns logs into structured records: order fills, market resolutions, position transfers. Normalization applies consistent units, timestamps, and a single outcome perspective across markets. Storage lands the result in parquet files or a time-series database sized for the query patterns you expect. Querying is the layer your bots, backtests, and dashboards actually touch.
A minimal build order looks like this:
Stand up an indexer against Polygon RPC or a bulk export CLI and checkpoint block heights.
Write decode logic for order-filled and resolution events, with retries on RPC timeouts.
Normalize outcome tokens, wallet addresses, and price units into a single schema.
Write to parquet or a columnar store, partitioned by market and date.
Expose a query layer for backtests, dashboards, or live bot signals.
Snapshot cadence matters more than most teams expect. Sampling the CLOB orderbook every few seconds versus every minute changes what microstructure questions you can answer later, so decide cadence before you backfill, not after.
Choosing your data source: on-chain, datasets, or APIs
Three practical routes exist for getting Polymarket history into a usable form, and each fits a different resource profile.
Direct Polygon RPC fetch: full control over raw events, but you own log parsing, reorg handling, and rate limit management indefinitely.
Community bulk datasets: the SII-WANGZJ/Polymarket_data project ships prebuilt orderfilled, trades, markets, quant, and users parquet files with CLI commands for fetching and decoding on-chain blocks, which saves the initial backfill work.
CLOB V2 SDKs: built primarily for order placement and execution, they expose live orderbook state but are not designed as a historical archive, so pairing them with a separate history source is standard.
Normalized data APIs: trade the flexibility of raw access for pre-cleaned tables, cutting integration time from weeks to hours.
A technical breakdown of Polymarket’s API layers notes that the platform’s data API covers user-level analytics, including full trade history with timestamps and sizes, settled on Polygon, which is useful context when deciding how much of the raw chain you actually need to touch yourself.
Pro Tip: Pull a small time slice from a community dataset before committing to a full backfill: it exposes schema quirks and gaps faster than reading documentation alone.

Why raw on-chain decoding breaks pipelines
Polymarket settles trades on Polygon, and the on-chain event log alone does not tell you which human decision produced a given fill. Several decoding steps trip up pipelines that look correct on the surface.
Condition ID resolution: each market’s condition ID must be mapped to human-readable metadata such as question text and resolution date, and that mapping lives outside the raw trade logs.
NegRisk contract noise: multi-outcome markets route trades through NegRisk wrapper contracts, and filtering out contract-level noise from genuine user trades takes deliberate rules, not a simple address check.
Token perspective: every market has YES and NO outcome tokens, and mixing perspectives without normalizing to one side produces inconsistent price and volume series across markets.
Maker and taker signs: fills need consistent signed volume conventions, or aggregated P&L and flow metrics silently invert.
Reorg protection: blocks can be reorganized shortly after being produced, so writes need a confirmation buffer before being treated as final.
None of these are exotic problems, but skipping any one of them produces a dataset that looks complete and is quietly wrong.
Schemas worth building: trades, wallets, markets, candles
Four tables cover most backtesting and research needs and getting their fields right up front saves rework later.
Table | Key fields | Purpose |
|---|---|---|
trades | timestamp, block, market_id, outcome_id, wallet, side, price, size, maker_taker, tx_hash | Trade-level fills for replay and P&L |
users | wallet, split_type, position_side, cumulative_volume | Wallet-level activity and position tracking |
quant | market_id, timestamp, yes_price, yes_size, signed_volume | Unified YES-perspective series for cross-market modeling |
ohlcv | market_id, interval_start, open, high, low, close, volume | Model inputs at fixed time intervals |
The SII-WANGZJ/Polymarket_data project’s quant.parquet and users.parquet files follow this general pattern, splitting wallet activity by position type and unifying prices to a single outcome perspective so you are not reconciling YES and NO series by hand. Pick a snapshot cadence for OHLCV candles that matches your shortest research question: a one-minute cadence supports microstructure work that a five-minute cadence cannot.
What deep history actually unlocks
Once trades, wallets, and candles are normalized, four workflows become straightforward.
Backtesting: replay trade-level fills in timestamp order, aligning each fill against the OHLCV snapshot in force at that moment to simulate realistic entry and exit prices.
Smart Money scoring: compute a rolling P&L per wallet, then weight it by trade frequency and position persistence, since a wallet with consistent, sized positions over time is a stronger signal than one large lucky bet.
ML feature generation: derive time-window aggregates, rolling volume, price volatility, and orderbook imbalance proxies as inputs for classification or ranking models.
Microstructure research: study spread behavior, order flow imbalance, and price impact around resolution events using dense CLOB snapshots.
Each of these depends on the same underlying normalization work described above. A guide to backtesting prediction market strategies walks through replay logic in more depth, and separate coverage of identifying Smart Money wallets on Polymarket details wallet scoring methods beyond the simple version above.
Keeping a pipeline trustworthy over time
A pipeline that works on day one can quietly drift once markets change shape or volume grows.
Write unit tests for decode logic against known historical blocks with expected output.
Require multiple block confirmations before treating a write as final, and add reorg detection on top.
Monitor daily event counts against a rolling baseline to catch silent gaps early.
Test new decode logic against replay data before pointing it at a live feed.
Keep a hot store for recent data your bots query constantly and a cold store for full history you query rarely.
Pro Tip: Log the raw event count alongside the decoded record count for every batch: a growing gap between the two numbers is usually the first sign of a decoding bug.
How Assymetrix removes the decoding burden
Assymetrix indexes Polymarket on-chain activity going back to September 2020, covering more than 900 million events and over 200 million OHLCV snapshots, all delivered as normalized tables rather than raw logs.
Polygon RPC management, condition ID mapping, and NegRisk filtering are handled externally.
Trade fills, wallet activity, and resolution events are delivered pre-normalized to a single outcome perspective.
API calls typically request a market’s trade history or candle series over a date range and return clean rows ready for a dataframe.
Documentation for Polymarket-specific endpoints lives at data.assymetrix.com/polymarket, alongside broader Python integration examples.
This does not replace the value of understanding the underlying decode logic covered above. It replaces the need to maintain it yourself in production.
Build versus buy: a practical decision checklist
The right answer depends on four questions, not on preference. How novel is your research question: does it need raw, unfiltered fields, or standard OHLCV and trade tables? How much fidelity do you need on token perspective and maker/taker signs? Does your team have bandwidth to maintain a decoder through Polymarket’s contract changes? And how fast do you need results, this week or this quarter? Institutional research also needs reproducibility: a documented, versioned dataset is easier to audit than a hand-rolled pipeline nobody remembers the details of a year later.
— Dean
Getting started with the Assymetrix Data API
Raw Polymarket history is scattered across on-chain logs, wrapper contracts, and shifting condition IDs, and building a reliable decoder is a real engineering project, not a weekend script. The Assymetrix Data API at data.assymetrix.com/api skips that work entirely, delivering normalized trade fills, wallet activity, resolution data, and OHLCV candles across more than 900 million indexed events dating back to September 2020.

Request an API key to pull historical trade and candle data directly into your environment.
Try the sandbox against a single market before committing to a full historical backfill.
Request a bulk export if your research needs the full history in one file rather than paginated calls.
Full documentation and endpoint references for Polymarket coverage live at data.assymetrix.com/polymarket, with the main product and sign-up flow at the Data API page.
Where to go next for datasets and docs
The SII-WANGZJ/Polymarket_data repository for bulk parquet files and CLI decoding tools.
The Chainstack developer guide for API layer and CLOB V2 details.
Assymetrix’s Polymarket analytics guide for the full on-chain to analysis-ready workflow.
Sources
FAQ
What does Polymarket historical data actually include?
It spans trade-level fills, CLOB orderbook snapshots, wallet-level activity, market resolution events, and OHLCV candles built from those fills. Most research projects need at least trades and candles, with wallet data added for Smart Money analysis.
Why is fetching raw Polymarket data directly so complex?
Raw data lives as on-chain logs on Polygon, and turning those into usable rows means resolving condition IDs to market metadata, filtering NegRisk contract noise from genuine trades, and normalizing YES and NO token perspectives consistently. Skipping any of these steps tends to produce data that looks right but is not.
Should I use the community Polymarket_data dataset or an API?
The SII-WANGZJ/Polymarket_data dataset is a strong starting point for teams that want raw parquet files and are comfortable maintaining their own decode logic going forward. A normalized API is the faster route when you want clean tables without owning that maintenance.
What can I build with deep Polymarket history?
Trade-level and candle history support strategy backtesting, Smart Money wallet scoring, machine learning feature generation, and market microstructure research. Each of these depends on having consistent, normalized fields rather than raw, unfiltered logs.
Does Assymetrix require blockchain decoding on my end?
No. The Assymetrix Data API delivers pre-normalized Polymarket trade fills, wallet activity, resolution data, and OHLCV candles across more than 900 million indexed on-chain events dating back to September 2020, so your team queries clean tables instead of parsing logs.
Other Blog



