{"id":"5e3becf1-de75-41b3-89b3-56138ff400eb","arxiv_id":"2605.30802","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Independent aggregation of LLMs reaches 83.43% accuracy on 1,189 KalshiBench questions, 1.01 points above the best single model, while deliberative consensus drops to 76% and error correlations limit further gains.","lead":"The paper evaluates multi-agent LLM systems for resolving outcomes in prediction markets, finding that independent aggregation slightly beats single models while group deliberation performs worse due to error spread. A smart generalist might read it to understand the practical limits and hybrid designs for AI decision systems in forecasting or betting platforms.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of the 1,189 KalshiBench questions and whether date-filtered Exa retrieval truly isolates reasoning quality","rationale":"The concern identified is identical to the reader's weakest assumption. Because the full methods, retrieval logs, and per-question breakdowns are not visible in the provided abstract, this remains the single most load-bearing unverified precondition for the central empirical claim. No other internal inconsistency or hidden assumption rises to the same level of risk for the reported accuracy ordering.","tokens_in":1766,"tokens_out":298,"duration_ms":16351,"concrete_test":"Partition the 1,189 questions into strata by market category and resolution horizon; recompute per-stratum accuracies and the aggregation gain; if the +1.01pp advantage disappears or reverses in any major stratum, the overall result does not generalize.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim (83.43% accuracy for confidence-weighted aggregation, +1.01pp over best single model) requires that the benchmark questions are representative of real prediction-market resolution difficulty and that sharing a date-filtered Exa evidence layer equalizes retrieval quality so that measured differences reflect only reasoning/aggregation. If the sample over-weights easy or short-horizon markets, or if retrieval still differs in coverage or relevance across models despite the filter, the small observed gain could be an artifact rather than a general property of multi-agent oracles.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript evaluates multi-agent LLM architectures for resolving outcomes in prediction markets using the KalshiBench dataset of 1,189 resolved questions. It compares single-LLM baselines (GPT-5 Nano, DeepSeek V3, Llama-3.3-70B) against independent aggregation with confidence-weighted voting and deliberative consensus, all using a shared date-filtered Exa evidence retrieval layer. The central claim is that independent aggregation achieves 83.43% accuracy, outperforming the best single model by 1.01 percentage points, while deliberative consensus performs worse at approximately 76% due to error propagation. Error correlations between 0.529 and 0.689 are reported as limiting ensemble gains below the Condorcet bound. The paper proposes hybrid routing criteria for auto-resolving unanimous high-confidence cases at 97.87% accuracy on 47% of the data, flagging the rest for human review.","tokens_in":1898,"tokens_out":576,"duration_ms":21656,"significance":"If the results hold after addressing methodological gaps, this work provides direct empirical measurements on an external benchmark (KalshiBench) showing modest benefits from simple confidence-weighted aggregation in multi-agent oracles while highlighting limits from correlated errors and the value of hybrid escalation. It contributes concrete data toward practical oracle system design without relying on self-referential derivations or fitted parameters.","major_comments":[{"comment":"Methods/Experiments section: The manuscript reports concrete accuracy numbers (83.43%, ~76%, 97.87%) and error correlations (0.529-0.689) but provides insufficient detail on implementation (prompting, confidence elicitation, aggregation mechanics), statistical significance testing for the 1.01pp gain, baseline configurations, and full error analysis. This absence is load-bearing for assessing the central empirical claim.","section":"Methods/Experiments section"},{"comment":"Results/Discussion on benchmark: The headline claim of 83.43% accuracy for confidence-weighted aggregation requires that the 1,189 KalshiBench questions form a representative sample of real-world prediction market resolution difficulty. The paper does not discuss selection criteria, horizon/difficulty distribution, or comparison to actual Kalshi markets, so the small observed gain could be an artifact of sample composition.","section":"Results/Discussion on benchmark"},{"comment":"Methods on retrieval: The claim that date-filtered Exa retrieval isolates reasoning quality from retrieval differences is load-bearing for attributing performance gaps to aggregation. No ablation on retrieval coverage, relevance, or model-specific utilization of the shared evidence layer is reported to validate the isolation.","section":"Methods on retrieval"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed feedback. We address each major comment below, indicating planned changes to strengthen the manuscript.","responses":[{"response":"We agree that greater implementation detail is required for reproducibility and rigorous evaluation. The revised manuscript will expand the Methods section with the exact prompts and output formats used for each model, the procedure for eliciting and normalizing confidence scores, the precise mechanics and weighting formula for independent aggregation, full baseline configurations, and an extended error analysis including per-question and per-category breakdowns. We will also add statistical significance testing (e.g., McNemar's test with bootstrap intervals) for the 1.01pp gain. These changes will be incorporated.","revision_made":"yes","referee_comment":"[Methods/Experiments section] Methods/Experiments section: The manuscript reports concrete accuracy numbers (83.43%, ~76%, 97.87%) and error correlations (0.529-0.689) but provides insufficient detail on implementation (prompting, confidence elicitation, aggregation mechanics), statistical significance testing for the 1.01pp gain, baseline configurations, and full error analysis. This absence is load-bearing for assessing the central empirical claim."},{"response":"KalshiBench is drawn directly from resolved Kalshi market questions, supplying an authentic sample of real-world resolution tasks. To address the concern, the revision will add a dedicated subsection describing the question selection criteria, the distribution of time horizons and market categories within the 1,189 questions, and available comparisons to performance on the broader Kalshi platform. This will allow readers to assess whether the modest gain generalizes beyond the current sample.","revision_made":"yes","referee_comment":"[Results/Discussion on benchmark] Results/Discussion on benchmark: The headline claim of 83.43% accuracy for confidence-weighted aggregation requires that the 1,189 KalshiBench questions form a representative sample of real-world prediction market resolution difficulty. The paper does not discuss selection criteria, horizon/difficulty distribution, or comparison to actual Kalshi markets, so the small observed gain could be an artifact of sample composition."},{"response":"The shared, date-filtered Exa layer ensures identical evidence is supplied to all models, which isolates differences to reasoning and aggregation. We acknowledge that explicit validation would strengthen the attribution. The revision will report retrieval coverage statistics, relevance indicators where available, and any observed model-specific patterns in evidence use. A complete model-specific ablation is constrained by the shared-layer design, but the added metrics will support the isolation claim.","revision_made":"partial","referee_comment":"[Methods on retrieval] Methods on retrieval: The claim that date-filtered Exa retrieval isolates reasoning quality from retrieval differences is load-bearing for attributing performance gaps to aggregation. No ablation on retrieval coverage, relevance, or model-specific utilization of the shared evidence layer is reported to validate the isolation."}],"tokens_in":1634,"tokens_out":627,"duration_ms":21902,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core result is that independent aggregation with confidence-weighted voting hits 83.43% accuracy on 1,189 resolved Kalshi questions, beating the best single model by 1.01 points. Deliberative consensus drops to around 76%, which the authors tie to error propagation. They also report error correlations between 0.529 and 0.689 across models and show that routing only unanimous high-confidence cases to auto-resolution gives 97.87% accuracy on 47% of the data.\n\nWhat stands out is the direct measurement of those correlations and the practical routing rule for hybrid AI-human systems. The setup uses a shared, date-filtered Exa evidence layer across agents, which at least tries to hold retrieval quality constant. The benchmark itself is new to this subfield.\n\nThe gains are small, and the paper does not claim otherwise. Without the full methods section it is hard to judge whether the 1,189 questions are representative of real market difficulty or whether the date filter truly equalizes retrieval across models. The drop in deliberative performance is reported but not explored in depth. No statistical significance tests or ablation details appear in the abstract.\n\nThis is useful for people already working on LLM oracles for prediction markets or similar resolution tasks. It is not a theoretical advance and does not introduce new algorithms. A serious referee could check the implementation details, the exact voting rule, and whether the benchmark sample is biased toward easier questions. I would send it to review rather than desk-reject, mainly because the empirical comparison is concrete and the hybrid-routing suggestion is testable.","headline":"This paper runs a straightforward empirical test of multi-agent LLM setups on a Kalshi prediction-market benchmark and finds that simple confidence-weighted aggregation edges out single models by about 1 point while deliberation hurts.","tokens_in":2356,"tokens_out":407,"would_cite":false,"duration_ms":9798,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Independent aggregation of multiple LLMs reaches 83.43 percent accuracy on resolved prediction market questions, 1.01 points above the strongest single model.","keywords":["multi-agent LLMs","prediction market oracles","confidence-weighted voting","KalshiBench","hybrid AI-human systems","error correlations","oracle resolution"],"falsifier":"Running the same independent-aggregation procedure on a new collection of resolved prediction-market questions and finding that it no longer exceeds the accuracy of the best single model would falsify the performance claim.","tokens_in":2671,"feed_emoji":"📊","tokens_out":721,"duration_ms":19667,"temperature":0.7,"pith_summary":"The paper tests whether multi-agent LLM systems can resolve uncertain events for prediction markets more reliably than any one model by itself. It runs both independent voting and group deliberation setups against single-model baselines on 1,189 real Kalshi questions, all using the same date-filtered evidence source. A reader would care because accurate automated resolution removes a major cost and delay barrier that currently limits how widely prediction markets can be used. The results show small gains from parallel aggregation but clear harm from debate, plus a practical way to hand off the hardest cases to humans.","feed_headline":"LLM aggregation hits 83.43% on market resolutions","feed_subtitle":"Independent voting edges out single models by one point while debate lowers accuracy, allowing hybrid human review to reach 97.87% on nearly","key_machinery":"Independent aggregation using confidence-weighted voting across multiple LLMs that share a common date-filtered evidence layer.","core_discovery":"Independent aggregation with confidence-weighted voting achieves 83.43 percent accuracy on the 1,189 KalshiBench questions, outperforming the best single LLM by 1.01 points, while deliberative consensus falls to roughly 76 percent because error propagation during debate allows confidently wrong agents to flip correct ones. Measured error correlations between 0.529 and 0.689 across models explain why ensemble gains stay well below the theoretical ceiling. The authors therefore propose a hybrid routing rule that auto-resolves only unanimous high-confidence questions at 97.87 percent accuracy for 47 percent of the dataset and escalates the rest to human review.","pith_inferences":["Disagreement among agents can function as an efficient triage signal for routing to humans in other high-stakes AI decision pipelines.","The results suggest that parallel independent reasoning may be preferable to interactive debate for any ensemble forecasting task where models share similar training data.","Increasing model diversity could lower the observed error correlations and expand the fraction of questions that can be auto-resolved without human review."],"forward_implications":["Deliberative consensus degrades accuracy below every single-model baseline because error propagation allows wrong agents to override correct ones.","Error correlations between 0.529 and 0.689 across models place a hard limit on how much any ensemble can improve results.","Many questions remain uncorrectable by any multi-agent architecture, so escalation to humans is required for those cases.","Auto-resolving only unanimous high-confidence questions delivers 97.87 percent accuracy on 47 percent of the dataset."],"fun_headline_variants":["Independent aggregation attains 83.43% accuracy on KalshiBench","Deliberative consensus achieves 76% accuracy on resolutions","Hybrid routing auto resolves 47% at 97.87% accuracy","Error correlations limit ensemble gains to 0.529-0.689 range"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That filtering retrieval by publication date fully removes differences in what each model knows, leaving only differences in how well each reasons.","fun_headline_variants_meta":{"raw":{"variants":["Independent aggregation attains 83.43% accuracy on KalshiBench","Deliberative consensus achieves 76% accuracy on resolutions","Hybrid routing auto resolves 47% at 97.87% accuracy","Error correlations limit ensemble gains to 0.529-0.689 range"]},"model":"grok-4.3","cost_usd":0.005911,"raw_usage":{"total_tokens":2867,"prompt_tokens":790,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":59112000,"prompt_tokens_details":{"text_tokens":790,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2003,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":790,"tokens_out":74,"duration_ms":14599,"temperature":1.0,"reasoning_tokens":2003,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T20:57:53.215092+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same independent-aggregation procedure on a new collection of resolved prediction-market questions and finding that it no longer exceeds the accuracy of the best single model would falsify the performance claim.","supporting_citations":[],"review_version":1}