{"id":"d1886b8d-61d5-4cb1-a6b0-e3338d3f9105","arxiv_id":"2507.08921","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper claims Polymarket betting prices outperformed traditional polling in predicting the 2024 U.S. presidential election, especially in swing states.","lead":"A new preprint compares Polymarket betting prices with polling averages for the 2024 U.S. presidential election and reports that the market predicted Donald Trump's victory more accurately than polls did. The study is a single-election case study whose central comparison may be flawed because it treats polling vote shares as win probabilities.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Polling average is plotted as 'Trump Win Probability' without conversion; the apparent Polymarket advantage may be an artifact of comparing vote share to a win-probability contract.","rationale":"I agree with the reader's weakest assumption: the paper compares a polling average, which estimates vote intention, to a Polymarket contract price, which estimates the probability of a binary election outcome. These are different estimands, and the paper does not convert one into the other. The figures label both as 'Trump Win Probability,' which makes the comparison seem direct, but it is not. In the predictive analysis, the BSTS model forecasts the next value of the poll average itself; the resulting predictive interval is an interval for the poll number, not for the election outcome. Interpreting the location of that interval relative to 0.5 as 'the polling data favor Harris' is therefore not a statement about predictive accuracy of polls for the election. The paper has real strengths: it is transparent about data sources, supplies code and supplementary materials, separates descriptive from predictive claims, and acknowledges limitations such as market manipulation and limited generalizability. Those strengths do not repair the invalid comparison. A secondary concern is the single-election, single-outcome design, but even a multi-election study would be undermined by comparing raw vote share to win probability, so the axis mismatch is the more load-bearing issue. The correct next step is a proper scoring-rule comparison in which polling data are first converted into win probabilities using a validated model. Until that is done, the central claim that Polymarket was superior to polling is not established, and the reader's REJECT verdict should stand unchanged.","tokens_in":19662,"tokens_out":2774,"duration_ms":38642,"concrete_test":"Re-do the comparison using a calibrated conversion from polling averages to win probabilities: for each day and state, take the polling margin and a historical national/state polling-error distribution (e.g., a FiveThirtyEight/Silver Bulletin style model) to compute P(Trump wins the Electoral College or that state); then compare Polymarket prices against the converted poll probabilities using Brier score and log score across all dates. If Polymarket no longer outperforms after conversion, the claimed superiority is an artifact; if it still outperforms, the claim survives. This test directly addresses the axis mismatch in Methods 2.1 and Figures 1–2.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on placing a daily polling average—a mean vote-intention estimate—on the same axis as a Polymarket contract price, which is a probability of winning the Electoral College. This occurs in Methods 2.1 and in Figures 1 and 2, where both series are labeled 'Trump Win Probability.' A poll showing 45% two-party vote share is not a 45% chance of winning; depending on the distribution of polling error, 45% vote share can correspond to a high or low win probability. By construction, a vote-share series will hover near 0.5 and often below it, while a market probability can range more freely. The BSTS forecasting section (Equations 5–6) then forecasts each series itself and treats the forecast of the polling average as a forecast of Trump's win probability, so the statement that 'almost all polling predictions favor Harris' (Section 3, Figure 4) is an artifact of the axis mismatch, not evidence about predictive skill. The authors acknowledge many limitations but do not address this comparability condition. Without converting poll margins into win probabilities via a validated model, the headline conclusion is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares daily Polymarket contract prices with daily polling averages from FiveThirtyEight for the 2024 U.S. presidential election, at the national level and in seven swing states, using descriptive time-series plots and Bayesian Structural Time Series (BSTS) models. The authors conclude that Polymarket was superior to polling in predicting the outcome, particularly in swing states, and suggest that betting markets could be used to predict elections and other events. The manuscript includes model fits, variable importance analysis, rolling forecasts, and an extensive limitations discussion, and states that code and data are available in supplementary materials.","tokens_in":19817,"tokens_out":4798,"duration_ms":55434,"significance":"If the comparison were valid, the paper would be a useful contribution to the ongoing debate on prediction markets versus traditional polls, adding a large-market, state-level case study with Bayesian uncertainty quantification. The authors also deserve credit for acknowledging market-manipulation concerns and for making code and data available. However, the central comparison is compromised by the fact that the two series measure different quantities, and the evaluative claims rest on informal visual criteria and a single election; in its current form the paper does not establish that betting markets are better predictors.","major_comments":[{"comment":"The polling data plotted as 'Trump Win Probability' are raw or averaged two-candidate vote shares (preference estimates), not probabilities of winning the Electoral College. A 45% vote share does not correspond to a 45% win probability; the mapping depends on the distribution of polling error and the correlation of state errors. Because the BSTS forecasting model in Equations (5)–(6) is applied directly to this vote-share series, the claim in Section 3 that 'almost all polling predictions favor Harris' is an artifact of comparing a vote-share quantity to a market-implied win probability, not evidence of inferior predictive skill. A valid comparison requires converting poll margins into win probabilities through a state-level forecast model with an explicit error structure before scoring.","section":"2.1, Figures 1–2, Figure 4"},{"comment":"The conclusion that Polymarket was 'superior' is based on visual inspection of whether posterior predictive intervals fall above 50% at selected horizons, not on any proper scoring rule (e.g., Brier score, logarithmic score, calibration test). With one election, seven states, and eight evaluation time points, no adjustment for multiple comparisons or selection is made, and no statistical significance test is reported. The paper should compare the two forecasts using a pre-specified proper scoring rule and an appropriate test of equal predictive ability, rather than interpreting plotted intervals after the outcome is known.","section":"3, Section 4.2, Figure 4"},{"comment":"The authors acknowledge that a single trader placed large pro-Trump bets across many accounts and that Polymarket was subject to wash-trading accusations during the exact period in which the market separated from the polls (October 2024). Given that a prediction-market price can be moved by capital as well as information, the headline claim of superiority requires robustness analysis: the comparison should at least be repeated excluding or downweighting the potentially manipulated post-October period, or with an adjustment for the large trader's positions.","section":"4.1, 4.3, 4.4"},{"comment":"The BSTS variable-importance analysis for the national market uses state-level Polymarket prices as regressors for the national Polymarket price. This is a within-market decomposition, not an independent validation of market predictive power, and it does not address whether market prices added information beyond polls. The discussion should not present this analysis as evidence of superiority over polling.","section":"2.4, Figure 3b"}],"minor_comments":[{"comment":"The text contains several typos: 'Hilary Clinton' should be 'Hillary Clinton', the reference 'Fried and and Harris' has a duplicated 'and', and 'real-word data' in Section 4.4 should be 'real-world data'.","section":"1.2, References"},{"comment":"Several equations and inline symbols are garbled: 'If 20, then t is constant' appears to have missing symbols for the variance parameters, and the Results section contains 'corroborate this notion (2)' and '(0 and ¿0)' where the estimated variance components are intended; these should be typeset correctly.","section":"2.4, Results"},{"comment":"The legends in Figures 1–4 include 'Standard Deviation' without explaining that it refers to the standard deviation of the polling average on that day; the captions should state this explicitly.","section":"Figures 1–4"},{"comment":"The manuscript states 'election day on November 4th, 2024' and Figure 4 labels 'November 04, 2024'; the 2024 U.S. presidential election was held on November 5, 2024, so the date is incorrect in both places.","section":"Discussion and Figure 4"},{"comment":"The list of pollsters is extensive, but no information is given about the aggregation method, weighting, or inclusion criteria for the FiveThirtyEight polling averages; a brief description would help readers assess the comparability of the polling series.","section":"Appendix A"}],"recommendation":"reject","confidential_remarks":"I agree with the reader's rejection. The paper's central problem is that the two compared quantities are incommensurable: the polling series is a vote-share estimate, while the market price is a win probability, and the plotted and forecast 'Trump Win Probability' labels effectively assume what needs to be shown. The absence of proper scoring rules and the reliance on a single election reinforce the verdict. These are load-bearing issues that cannot be fixed by local edits; a substantially reworked comparison and evaluation design would be needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the bottom line: the paper is a clean, readable case study of Polymarket vs. polling for the 2024 U.S. election, but its central claim is invalid as stated because it compares a market price (a probability of winning the Electoral College) to a raw polling average (a vote-intention share) and labels both 'Trump Win Probability.' That mismatch is load-bearing: it explains most of the apparent Polymarket advantage.\n\nWhat's actually new and worthwhile: it's the first systematic look at Polymarket state-level data, with daily data and code in the supplement. The BSTS models are competently described and the authors are honest about limitations—single election, market manipulation, and the demographic skew of crypto bettors. The descriptive event-reactivity analysis (the post-assassination-attempt jump, the Harris-nomination dip) is interesting and well presented.\n\nThe soft spot is not minor. In Figures 1 and 2, and in the BSTS forecasts of Section 3, a polling average that hovers around 45% is treated as a 45% chance of winning. That's a category error. A 45% two-party vote share can correspond to a wide range of win probabilities depending on the distribution of polling error. The statement that 'almost all polling predictions favor Harris' is an artifact of that mismatch, not evidence about predictive skill. The paper never converts poll margins into win probabilities via a validated model, so the comparison is between different quantities. There's also no proper scoring rule and no formal statistical test; the 'superiority' rests on one election and visual inspection of overlapped intervals.\n\nThe authors acknowledge many limitations but never address this comparability condition. I also note that the market's late-October peak coincides with the large trader discussed in the text, so the same manipulation concern that the authors raise undermines the market's own accuracy claim.\n\nBottom line: this is a useful descriptive case study, but the headline conclusion is unsupported. I would not send it to peer review as is; the right next step is a revision that converts polls to win probabilities with a defensible error model, evaluated across multiple elections with proper scoring rules. If that were done, the empirical result might be worth a short note. As is, I'd desk reject.","headline":"A well-written case study undone by comparing raw polling averages to market win probabilities on the same axis; the claimed Polymarket superiority is largely an artifact.","tokens_in":20403,"tokens_out":3664,"would_cite":false,"duration_ms":39622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Polymarket's daily betting prices predicted the 2024 U.S. presidential election outcome better than traditional polling, nationally and in most swing states.","keywords":["prediction markets","Polymarket","election forecasting","polling accuracy","Bayesian structural time series","wisdom of crowds","2024 presidential election","swing states"],"falsifier":"Re-run the comparison on the 2016 or 2020 elections using the same daily market and polling series and the same BSTS forecasting protocol: if the market's mid-October lead over polling does not reproduce in either election, or disappears once both sources are converted to a common outcome space, the 2024 result alone is too weak to carry the paper's general conclusion.","tokens_in":19414,"feed_emoji":"🗳️","tokens_out":6630,"duration_ms":72166,"temperature":0.7,"pith_summary":"Political elections are hard to forecast because polls ask whom people intend to vote for, not whom they expect to win. This paper argues that Polymarket, the largest cryptocurrency-based betting market, captured the latter and, in the 2024 presidential election, did a better job than traditional polling. The authors compare a daily Trump win probability from Polymarket contract prices with the aggregated FiveThirtyEight polling average, first visually and then with Bayesian structural time series forecasts. The market favored Trump at nearly every time point, reacted sharply to the assassination attempt, Harris's entry, and the debates, and by mid-October had forecast intervals that stayed above the 50 percent line. Polling averages never put Trump ahead and on Election Day still pointed to Harris, so the paper concludes that market-based forecasts are a viable, possibly superior, complement to polls.","feed_headline":"Polymarket beat polling on 2024 election calls","feed_subtitle":"Daily bet prices pointed to a Trump win by mid-October; poll averages never did.","key_machinery":"The analysis runs through a date-aligned pair of daily time series: Polymarket's closing price on the Trump full-ballot contract, treated as an implied win probability, and the average of all available polls from FiveThirtyEight, treated the same way. The predictive engine is a Bayesian structural time series model with a local-level state equation, fitted with spike-and-slab priors on regressors, which identifies Pennsylvania and Michigan as the state markets driving the national market and produces rolling forecasts with 95 percent predictive intervals. The model's fitted variance terms let the paper describe the market series as a random walk and the polling series as noise around a fixed mean, which is the formal reason the market forecasts are wider but more accurate.","core_discovery":"The central claim is that Polymarket was superior to polling in predicting the 2024 presidential election, both nationally and in five of the seven swing states, with Michigan and Wisconsin too close to call in either source. The market's daily contract price, read as the probability Donald Trump would win the presidency, called the winner early and stayed on the right side of the 50 percent decision boundary from mid-October onward. The polling average, built from preference questions, not only failed to call the result but got the direction wrong on Election Day. The paper presents this as the first systematic comparison of Polymarket data against traditional polls, and reads the result as support for the idea that a crowd wagering money can aggregate expectations more accurately than a crowd answering survey questions.","pith_inferences":["The strongest hidden comparison would be to map poll vote shares into win probabilities with an electoral-college model; on that common scale, the market's edge may shrink, and the paper does not run that test.","Michigan and Wisconsin, where the market added no clear signal, suggest the market's value may be limited to races with an identifiable structural lean, not genuinely tied states.","Adding a direct 'who do you think will win' question to standard polls could isolate whether the market's advantage is expectations rather than money, without relying on a crypto platform that was legally restricted inside the United States.","A single election is a sample of one; before recommending betting markets, the same protocol should be run across many political and non-political binary events to see if the superiority generalizes."],"forward_implications":["Campaigns and media could treat daily market prices as an early-warning signal for where resources should go, since the market saw Georgia and North Carolina as unwinnable for Harris months before the polls did.","Forecasters could use market data to measure how events move the race hour by hour, a granularity most polls cannot offer.","If markets are adopted as forecasting tools, regulators and platforms would need identity checks and anti-manipulation rules, because a single large bettor or wash trades can shift prices.","The 2024 election becomes a test case for wisdom-of-crowds theory, turning a stochastic market into a calibration point for collective judgment."],"supporting_citations":[{"why":"Supplies the daily Polymarket state and national closing prices used as the market probability series.","marker":"(Andrade, 2024)"},{"why":"Provides the BSTS model specification and spike-and-slab variable selection used for fitting and forecasting.","marker":"(Brodersen et al., 2015)"},{"why":"Gives the local-level structural time series equations the paper's predictive models are built on.","marker":"(Scott and Varian, 2013)"},{"why":"Establishes the prior question of whether markets outperform polls as election predictors, which this study extends to Polymarket.","marker":"(Erikson and Wlezien, 2008)"},{"why":"Supplies the economic rationale for reading prediction market prices as probabilities.","marker":"(Wolfers and Zitzewitz, 2004)"},{"why":"Provides the wisdom-of-crowds theory the paper uses to explain why a betting crowd could beat survey respondents.","marker":"(Surowiecki, 2004)"},{"why":"Documents the large pro-Trump bets across multiple accounts that the paper cites as the main manipulation threat to the market signal.","marker":"(Osipovich, 2024)"},{"why":"Records the settlement restricting Polymarket to users outside the United States, a key caveat about who the market's crowd actually is.","marker":"(CFTC, 2022)"}],"fun_headline_variants":["Betting markets beat polls in predicting elections","Polymarket outperformed traditional polling in 2024","Money-wagering crowds forecast elections better than polls","Polymarket's daily prices called 2024 election early","Wisdom of crowds wins over polling in election tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that a poll's vote-preference average can be plotted on the same 'Trump win probability' axis as a Polymarket contract that actually pays out on who wins the Electoral College, and if those quantities measure different things, the market's apparent advantage is an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Betting markets beat polls in predicting elections","Polymarket outperformed traditional polling in 2024","Money-wagering crowds forecast elections better than polls","Polymarket's daily prices called 2024 election early","Wisdom of crowds wins over polling in election tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000653,"raw_usage":{"total_tokens":3005,"prompt_tokens":968,"completion_tokens":2037,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":1960}},"tokens_in":584,"tokens_out":2037,"duration_ms":18035,"temperature":1.0,"reasoning_tokens":1960,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:10:04.757045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison on the 2016 or 2020 elections using the same daily market and polling series and the same BSTS forecasting protocol: if the market's mid-October lead over polling does not reproduce in either election, or disappears once both sources are converted to a common outcome space, the 2024 result alone is too weak to carry the paper's general conclusion.","supporting_citations":[{"cited_title":"The Wisdom of Crowds: Why the Many Are Smarter than the Few and How Collective Wisdom Shapes Business, Economies, Societies, and Nations","cited_arxiv_id":null,"evidence_quote":"Provides the wisdom-of-crowds theory the paper uses to explain why a betting crowd could beat survey respondents."},{"cited_title":"A mystery \\ 30 million wave of pro-trump bets has moved a popular prediction market, 2024","cited_arxiv_id":null,"evidence_quote":"Documents the large pro-Trump bets across multiple accounts that the paper cites as the main manipulation threat to the market signal."},{"cited_title":"Cftc orders event-based binary options markets operator to pay \\ 1.4 million penalty","cited_arxiv_id":null,"evidence_quote":"Records the settlement restricting Polymarket to users outside the United States, a key caveat about who the market's crowd actually is."}],"review_version":1}