Assymetrix Launches the Deepest Independent Prediction Market Data APIs
Read more
Read more
Assymetrix Launches the Deepest Independent Prediction Market Data APIs
Read more
Read more
Developers & Quants: Don't Use Polymarket Feed's 59% Sign in Order Flow
Developers & Quants: Don't Use Polymarket Feed's 59% Sign in Order Flow
Developers & Quants: Don't Use Polymarket Feed's 59% Sign in Order Flow
Developers and quants: why Polymarket feed direction matches on chain only 59% and how to build validated order flow pipelines using on chain joins and...

Developers & Quants: Don’t Use Polymarket Feed’s 59% Sign in Order Flow
For Polymarket microstructure and direction-sensitive features, join on-chain OrderFilled events rather than relying on the public WebSocket feed. Research measuring both against each other found feed-inferred direction matches on-chain ground truth only about 59% of the time, which is close enough to a coin flip to flip the sign of any direction-dependent measure you build. Always join on-chain fills before computing buy or sell pressure, effective spread, or price-impact estimators.
TL;DR:
Feed-inferred trade direction only matches on-chain truth about 59% of the time, making it unreliable for direction-sensitive metrics like spreads and impact estimates.
On-chain OrderFilled events are the authoritative source for trade aggressor sign, and joining them with feed data is essential for accurate analysis.
Accurate model building requires handling data quirks such as fragmented fills, token minting/burning, clock offsets, and liquidity normalization to avoid bias.
Smart Money wallets should be scored across multiple markets before indicating informed trading, and lead-lag signals must account for timing and settlement differences between venues.
Using pre-joined trade-level data via the Assymetrix Data API simplifies pipeline complexity and enhances reliability for live trading signals and backtesting.
AssymetrixBuild On Reliable Market DataAssymetrix unifies cross venue prediction market data, including pre joined trade level data, through one integration for research and trading systems.Explore Assymetrix
Table of Contents
Polymarket’s raw data layers: feed, REST, and on-chain fills
What the 59% agreement number means for your models
Order flow signals worth building: smart money, imbalance, spikes, divergence
Building the indicator: joins, features, and validation
Data engineering hazards that quietly bias your results
Replaying and backtesting Polymarket order-flow strategies
What holds up under replication and what stays exploratory
Assymetrix data API: one integration for Polymarket trade-level joins
Key research and replication packages behind this guide
Sources
Polymarket’s raw data layers: feed, REST, and on-chain fills
Polymarket exposes three distinct data layers, and each serves a different purpose. The WebSocket feed streams price_change and book_snapshot messages, which are useful for tracking quote dynamics and reconstructing depth at a point in time, but they carry no verified aggressor information. The CLOB REST API mirrors this same view for polling clients. The on-chain OrderFilled event, emitted by the Conditional Token Framework (CTF) Exchange, is the only layer that records the actual settlement: makerAssetId and takerAssetId together encode which side initiated the trade, making it the authoritative source for aggressor sign.
Building a correct join requires a few supporting pieces:
A versioned token ID to market slug map, since Polymarket’s outcome tokens are identified by long numeric IDs that shift across market instances
Market metadata (resolution date, question text, outcome labels) to attach context to raw fills
A stable schema for both feed and on-chain records so joins do not break across schema revisions
Treat the feed as a quote-dynamics source and the chain as a trade-truth source. Conflating the two is where most Polymarket order flow analysis breaks down.
What the 59% agreement number means for your models
The clearest empirical result in current Polymarket microstructure research is the sign-agreement rate between feed-inferred direction and on-chain ground truth, measured across a large joined archive of feed and on-chain events spanning calibration windows in 2026.

59% sign agreement between feed-inferred and on-chain direction on a large joined event archive, reported in recent microstructure research, means roughly four in ten trades get the wrong sign when you infer direction from the feed alone.
That error rate does not stay contained to one metric. It propagates:
Effective spread estimates widen or narrow depending on which trades get mis-signed, distorting cost-of-trading measures
Price-impact regressions attribute impact to the wrong side, weakening or inverting coefficients
Any buy-versus-sell pressure indicator built on feed-inferred sign inherits the noise directly, since the underlying label is wrong close to half the time
Direction-dependent metrics, in short, require on-chain sourcing as a baseline requirement, not an enhancement.
Order flow signals worth building: smart money, imbalance, spikes, divergence
Four signal families cover most of what a Polymarket order-flow indicator needs to capture.
Smart Money wallet entries. Track wallets by realized accuracy and position sizing over time, then score persistence: a wallet that shows up early and directionally correct across multiple unrelated markets is a stronger signal than one good call.
Depth-weighted orderbook imbalance. Compute imbalance at the top few price levels, then normalize by time-to-resolution, since imbalance near a market’s close carries different weight than imbalance weeks out.
Sudden volume spikes. Run a rolling z-score on trade volume per market and flag windows several standard deviations above baseline, tuning the window length to the market’s typical resolution horizon.
Cross-venue divergence. Pair Polymarket markets against equivalent Kalshi contracts on the same event, then measure lead-lag using a fixed comparison window rather than raw timestamp differences.
Pro Tip: Score wallet persistence across at least three independent markets before treating an address as a Smart Money signal, since single-market accuracy is often noise.
Cross-venue pairing is the hardest of the four to get right. Kalshi and Polymarket settle differently, quote at different granularities, and rarely share a clock, so lead-lag windows need to absorb timing slack without erasing the causal ordering you are trying to measure.
Building the indicator: joins, features, and validation
A production-grade order-flow indicator follows a fixed sequence from raw events to a validated signal.
Ingest both the WebSocket feed and on-chain OrderFilled events into columnar storage, Parquet works well, keeping raw records untouched for replay.
Map token IDs to market slugs through a versioned mapping table, since Polymarket reissues token IDs across market instances.
Join feed records to on-chain fills using price and a tight time window, always deferring to the on-chain record for sign.
Construct features: match-normalized signed notional, fragment-aware trade counts, and depth-weighted imbalance scaled by top-of-book liquidity.
Validate before trusting the output.
Validation is not optional given how easy these pipelines are to get subtly wrong:
Run a per-day sign-agreement check between your joined output and a held-out on-chain sample
Check spread-change concordance around large signed trades to confirm the direction label behaves as expected
Replay a single trading day end to end before scaling to a full historical backfill
A developer guide to fetching and parsing Polymarket’s orderbook covers the ingestion mechanics in more depth if you are starting the pipeline from scratch.
Data engineering hazards that quietly bias your results
Several structural quirks in Polymarket’s settlement data will bias a naive pipeline if left unhandled.
Fragmentation inflates counts. A single aggressive order can generate multiple maker-side OrderFilled records, so counting records as independent trades overstates maker-side activity. Group fragments by aggressive order and normalize by match, as recent identification-limits research recommends.
Mint and burn semantics break naive labeling. Some OrderFilled records represent token minting or burning rather than a simple buy or sell, and treating them as ordinary trades corrupts directional counts.
Clock offsets distort pairing. A single-vantage collector will show a constant offset relative to on-chain time, so pairing windows need to absorb that offset without collapsing genuine lead-lag structure.
Normalization needs a liquidity anchor. Per-market z-scoring and liquidity-aware scaling (dividing by top-10 depth) make features comparable across markets with very different trading volume.
Pro Tip: Run a clock-offset diagnostic against a handful of known large trades before trusting any lead-lag measurement across venues or between feed and chain.
Replaying and backtesting Polymarket order-flow strategies
A reproducible replay corpus needs raw feed ticks, on-chain fills, and orderbook snapshots stored together with enough history to span multiple market cycles, not just a single resolution window.
Split walk-forward, never randomly, since prediction markets carry strong time dependence toward resolution.
Model fees and slippage explicitly, and compare against a calibrated mid-price prior as your baseline. A synchronized Polymarket-Binance study found that raw multivariate microstructure features often fail to beat that mid prior out-of-sample, which makes it a demanding baseline rather than a token comparison.
Run null experiments and permutation tests on your lead-lag windows to confirm any measured edge survives when the causal ordering is scrambled.
Skipping the null-experiment step is the single most common way backtested Polymarket signals turn out to be artifacts of pairing-window choice rather than real structure.
What holds up under replication and what stays exploratory
Wallet persistence and cross-venue lead signals have shown up consistently across independent replication attempts, which puts them closer to operational than exploratory. Fine-grained microstructure features derived from raw feed data are more fragile: they are sensitive to pairing-window choice and often fail to beat a calibrated mid prior once tested out-of-sample, as the Polymarket-Binance corpus demonstrates.
Latency, storage cost, and model robustness trade against each other constantly in this data. Cite dataset releases and replication-package hashes explicitly in any research output. A replication package that joins feed archives to on-chain fills is a useful reference point for the kind of provenance a rigorous pipeline should preserve.
— Dean
Assymetrix data API: one integration for Polymarket trade-level joins
Building the pipeline described above from scratch means standing up on-chain indexing, feed capture, canonical ID mapping, and Smart Money scoring separately, then keeping all four in sync as Polymarket’s schema evolves. A single integration can provide trade-level joins between feed and on-chain fills, canonical token ID maps, wallet tracking, and historical indexes covering multiple prediction market venues.

Such a data feed supports use cases like bot ingestion pipelines that need clean signed trade data without running an on-chain indexer, AI agents consuming structured market data live, and backtests requiring reproducible historical depth rather than reconstructed data.
Full endpoint documentation lives at the Data API landing page, and Polymarket-specific developer resources, including on-chain join examples, are available through Polymarket data analysis for quants and developers. Start there to get trade-level granularity and live orderbook depth through one connection instead of four.

Key research and replication packages behind this guide
The Polymarket microstructure study, the fill-attribution research, and the Binance-paired corpus supply the evidence above. A public replication package and BitPulse’s market-data resources offer further tooling for readers building their own pipeline.
FAQ
What does order flow analysis reveal on Polymarket?
It reveals buy versus sell pressure, large wallet entries relative to news timing, and orderbook imbalance at specific probability levels, but only when direction is sourced from on-chain OrderFilled events rather than the public feed. Feed-inferred direction matches on-chain truth only about 59% of the time, so any signal built on feed-only sign is unreliable.
Why is Polymarket’s on-chain data more transparent than traditional markets?
Polymarket settles every trade through a public Conditional Token Framework (CTF) Exchange contract, so the full fill history is auditable on-chain rather than held privately by an exchange operator. That transparency lets researchers verify aggressor direction directly instead of inferring it from a proprietary feed, which is not possible on most traditional venues.
How do I detect Smart Money wallets on Polymarket?
Score wallets by tracking directional accuracy and position sizing across multiple independent markets rather than judging a single trade. A wallet that repeatedly enters early and correctly across unrelated markets is a stronger candidate than one large, isolated position.
How accurate is feed-inferred trade direction compared to on-chain data?
Feed-inferred direction agrees with on-chain ground truth in about 59% of trades based on a large joined event archive, which is far too noisy for direction-sensitive metrics like effective spread or price impact. On-chain OrderFilled events should be the direction source whenever a measure depends on trade sign.
Can I build a Polymarket order-flow indicator without running my own blockchain indexer?
Yes: the Assymetrix Data API provides pre-joined trade-level data and live orderbook depth for Polymarket through a single integration, removing the need to index on-chain events separately. Developer resources for Polymarket-specific joins are available at Polymarket data analysis for quants and developers.
Developers & Quants: Don’t Use Polymarket Feed’s 59% Sign in Order Flow
For Polymarket microstructure and direction-sensitive features, join on-chain OrderFilled events rather than relying on the public WebSocket feed. Research measuring both against each other found feed-inferred direction matches on-chain ground truth only about 59% of the time, which is close enough to a coin flip to flip the sign of any direction-dependent measure you build. Always join on-chain fills before computing buy or sell pressure, effective spread, or price-impact estimators.
TL;DR:
Feed-inferred trade direction only matches on-chain truth about 59% of the time, making it unreliable for direction-sensitive metrics like spreads and impact estimates.
On-chain OrderFilled events are the authoritative source for trade aggressor sign, and joining them with feed data is essential for accurate analysis.
Accurate model building requires handling data quirks such as fragmented fills, token minting/burning, clock offsets, and liquidity normalization to avoid bias.
Smart Money wallets should be scored across multiple markets before indicating informed trading, and lead-lag signals must account for timing and settlement differences between venues.
Using pre-joined trade-level data via the Assymetrix Data API simplifies pipeline complexity and enhances reliability for live trading signals and backtesting.
AssymetrixBuild On Reliable Market DataAssymetrix unifies cross venue prediction market data, including pre joined trade level data, through one integration for research and trading systems.Explore Assymetrix
Table of Contents
Polymarket’s raw data layers: feed, REST, and on-chain fills
What the 59% agreement number means for your models
Order flow signals worth building: smart money, imbalance, spikes, divergence
Building the indicator: joins, features, and validation
Data engineering hazards that quietly bias your results
Replaying and backtesting Polymarket order-flow strategies
What holds up under replication and what stays exploratory
Assymetrix data API: one integration for Polymarket trade-level joins
Key research and replication packages behind this guide
Sources
Polymarket’s raw data layers: feed, REST, and on-chain fills
Polymarket exposes three distinct data layers, and each serves a different purpose. The WebSocket feed streams price_change and book_snapshot messages, which are useful for tracking quote dynamics and reconstructing depth at a point in time, but they carry no verified aggressor information. The CLOB REST API mirrors this same view for polling clients. The on-chain OrderFilled event, emitted by the Conditional Token Framework (CTF) Exchange, is the only layer that records the actual settlement: makerAssetId and takerAssetId together encode which side initiated the trade, making it the authoritative source for aggressor sign.
Building a correct join requires a few supporting pieces:
A versioned token ID to market slug map, since Polymarket’s outcome tokens are identified by long numeric IDs that shift across market instances
Market metadata (resolution date, question text, outcome labels) to attach context to raw fills
A stable schema for both feed and on-chain records so joins do not break across schema revisions
Treat the feed as a quote-dynamics source and the chain as a trade-truth source. Conflating the two is where most Polymarket order flow analysis breaks down.
What the 59% agreement number means for your models
The clearest empirical result in current Polymarket microstructure research is the sign-agreement rate between feed-inferred direction and on-chain ground truth, measured across a large joined archive of feed and on-chain events spanning calibration windows in 2026.

59% sign agreement between feed-inferred and on-chain direction on a large joined event archive, reported in recent microstructure research, means roughly four in ten trades get the wrong sign when you infer direction from the feed alone.
That error rate does not stay contained to one metric. It propagates:
Effective spread estimates widen or narrow depending on which trades get mis-signed, distorting cost-of-trading measures
Price-impact regressions attribute impact to the wrong side, weakening or inverting coefficients
Any buy-versus-sell pressure indicator built on feed-inferred sign inherits the noise directly, since the underlying label is wrong close to half the time
Direction-dependent metrics, in short, require on-chain sourcing as a baseline requirement, not an enhancement.
Order flow signals worth building: smart money, imbalance, spikes, divergence
Four signal families cover most of what a Polymarket order-flow indicator needs to capture.
Smart Money wallet entries. Track wallets by realized accuracy and position sizing over time, then score persistence: a wallet that shows up early and directionally correct across multiple unrelated markets is a stronger signal than one good call.
Depth-weighted orderbook imbalance. Compute imbalance at the top few price levels, then normalize by time-to-resolution, since imbalance near a market’s close carries different weight than imbalance weeks out.
Sudden volume spikes. Run a rolling z-score on trade volume per market and flag windows several standard deviations above baseline, tuning the window length to the market’s typical resolution horizon.
Cross-venue divergence. Pair Polymarket markets against equivalent Kalshi contracts on the same event, then measure lead-lag using a fixed comparison window rather than raw timestamp differences.
Pro Tip: Score wallet persistence across at least three independent markets before treating an address as a Smart Money signal, since single-market accuracy is often noise.
Cross-venue pairing is the hardest of the four to get right. Kalshi and Polymarket settle differently, quote at different granularities, and rarely share a clock, so lead-lag windows need to absorb timing slack without erasing the causal ordering you are trying to measure.
Building the indicator: joins, features, and validation
A production-grade order-flow indicator follows a fixed sequence from raw events to a validated signal.
Ingest both the WebSocket feed and on-chain OrderFilled events into columnar storage, Parquet works well, keeping raw records untouched for replay.
Map token IDs to market slugs through a versioned mapping table, since Polymarket reissues token IDs across market instances.
Join feed records to on-chain fills using price and a tight time window, always deferring to the on-chain record for sign.
Construct features: match-normalized signed notional, fragment-aware trade counts, and depth-weighted imbalance scaled by top-of-book liquidity.
Validate before trusting the output.
Validation is not optional given how easy these pipelines are to get subtly wrong:
Run a per-day sign-agreement check between your joined output and a held-out on-chain sample
Check spread-change concordance around large signed trades to confirm the direction label behaves as expected
Replay a single trading day end to end before scaling to a full historical backfill
A developer guide to fetching and parsing Polymarket’s orderbook covers the ingestion mechanics in more depth if you are starting the pipeline from scratch.
Data engineering hazards that quietly bias your results
Several structural quirks in Polymarket’s settlement data will bias a naive pipeline if left unhandled.
Fragmentation inflates counts. A single aggressive order can generate multiple maker-side OrderFilled records, so counting records as independent trades overstates maker-side activity. Group fragments by aggressive order and normalize by match, as recent identification-limits research recommends.
Mint and burn semantics break naive labeling. Some OrderFilled records represent token minting or burning rather than a simple buy or sell, and treating them as ordinary trades corrupts directional counts.
Clock offsets distort pairing. A single-vantage collector will show a constant offset relative to on-chain time, so pairing windows need to absorb that offset without collapsing genuine lead-lag structure.
Normalization needs a liquidity anchor. Per-market z-scoring and liquidity-aware scaling (dividing by top-10 depth) make features comparable across markets with very different trading volume.
Pro Tip: Run a clock-offset diagnostic against a handful of known large trades before trusting any lead-lag measurement across venues or between feed and chain.
Replaying and backtesting Polymarket order-flow strategies
A reproducible replay corpus needs raw feed ticks, on-chain fills, and orderbook snapshots stored together with enough history to span multiple market cycles, not just a single resolution window.
Split walk-forward, never randomly, since prediction markets carry strong time dependence toward resolution.
Model fees and slippage explicitly, and compare against a calibrated mid-price prior as your baseline. A synchronized Polymarket-Binance study found that raw multivariate microstructure features often fail to beat that mid prior out-of-sample, which makes it a demanding baseline rather than a token comparison.
Run null experiments and permutation tests on your lead-lag windows to confirm any measured edge survives when the causal ordering is scrambled.
Skipping the null-experiment step is the single most common way backtested Polymarket signals turn out to be artifacts of pairing-window choice rather than real structure.
What holds up under replication and what stays exploratory
Wallet persistence and cross-venue lead signals have shown up consistently across independent replication attempts, which puts them closer to operational than exploratory. Fine-grained microstructure features derived from raw feed data are more fragile: they are sensitive to pairing-window choice and often fail to beat a calibrated mid prior once tested out-of-sample, as the Polymarket-Binance corpus demonstrates.
Latency, storage cost, and model robustness trade against each other constantly in this data. Cite dataset releases and replication-package hashes explicitly in any research output. A replication package that joins feed archives to on-chain fills is a useful reference point for the kind of provenance a rigorous pipeline should preserve.
— Dean
Assymetrix data API: one integration for Polymarket trade-level joins
Building the pipeline described above from scratch means standing up on-chain indexing, feed capture, canonical ID mapping, and Smart Money scoring separately, then keeping all four in sync as Polymarket’s schema evolves. A single integration can provide trade-level joins between feed and on-chain fills, canonical token ID maps, wallet tracking, and historical indexes covering multiple prediction market venues.

Such a data feed supports use cases like bot ingestion pipelines that need clean signed trade data without running an on-chain indexer, AI agents consuming structured market data live, and backtests requiring reproducible historical depth rather than reconstructed data.
Full endpoint documentation lives at the Data API landing page, and Polymarket-specific developer resources, including on-chain join examples, are available through Polymarket data analysis for quants and developers. Start there to get trade-level granularity and live orderbook depth through one connection instead of four.

Key research and replication packages behind this guide
The Polymarket microstructure study, the fill-attribution research, and the Binance-paired corpus supply the evidence above. A public replication package and BitPulse’s market-data resources offer further tooling for readers building their own pipeline.
FAQ
What does order flow analysis reveal on Polymarket?
It reveals buy versus sell pressure, large wallet entries relative to news timing, and orderbook imbalance at specific probability levels, but only when direction is sourced from on-chain OrderFilled events rather than the public feed. Feed-inferred direction matches on-chain truth only about 59% of the time, so any signal built on feed-only sign is unreliable.
Why is Polymarket’s on-chain data more transparent than traditional markets?
Polymarket settles every trade through a public Conditional Token Framework (CTF) Exchange contract, so the full fill history is auditable on-chain rather than held privately by an exchange operator. That transparency lets researchers verify aggressor direction directly instead of inferring it from a proprietary feed, which is not possible on most traditional venues.
How do I detect Smart Money wallets on Polymarket?
Score wallets by tracking directional accuracy and position sizing across multiple independent markets rather than judging a single trade. A wallet that repeatedly enters early and correctly across unrelated markets is a stronger candidate than one large, isolated position.
How accurate is feed-inferred trade direction compared to on-chain data?
Feed-inferred direction agrees with on-chain ground truth in about 59% of trades based on a large joined event archive, which is far too noisy for direction-sensitive metrics like effective spread or price impact. On-chain OrderFilled events should be the direction source whenever a measure depends on trade sign.
Can I build a Polymarket order-flow indicator without running my own blockchain indexer?
Yes: the Assymetrix Data API provides pre-joined trade-level data and live orderbook depth for Polymarket through a single integration, removing the need to index on-chain events separately. Developer resources for Polymarket-specific joins are available at Polymarket data analysis for quants and developers.
Other Blog



