{"id":"bfe5eb23-c102-416a-b92f-fb1c3c1f4789","arxiv_id":"2605.28255","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a competitive QA game, humans under-rely on correct AI suggestions 3.9% of the time and over-rely on incorrect ones 1.7% of the time, driven by confirmation bias and near-chance AI confidence when answers disagree.","lead":"The study ran a question-answering game where humans could delegate to AI without seeing its answer or adopt its suggestions after seeing them. Smart generalists should read it because the quantified patterns of under- and over-reliance point to concrete design changes that could make human-AI teams more reliable in practice.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Generalizability of observed suboptimal reliance rates from this competitive expert QA game to other human-AI contexts is untested","rationale":"The reader's weakest_assumption matches the load-bearing concern exactly. The UNVERDICTED verdict is appropriate because the generalization step is required for the strongest_claim to have force beyond the specific 24-match corpus, and no evidence for it is supplied.","tokens_in":1761,"tokens_out":312,"duration_ms":20826,"concrete_test":"Run an otherwise identical protocol but remove the competitive scoring and team-vs-team element (present the same questions to the same humans and AI agents in a solitary accuracy task); recompute the under- and over-reliance percentages on the new 1440+ adoption decisions. A shift >2 percentage points falsifies the implicit claim that the observed rates are not artifacts of the game format.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the 3.9% under-reliance and 1.7% over-reliance figures being evidence of generally suboptimal human decisions. These are measured in 24 matches involving expert humans, a competitive scoring game, and 16 specific AI agents. Nothing in the reported design tests whether the same rates (or even the same direction of bias) appear outside competition, with non-experts, or on different tasks. Without that, the leap from \"in this game humans sometimes miss correct AI output\" to \"humans make suboptimal collaboration decisions\" is the weakest link.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reports an observational study of human-AI collaboration in a competitive question-answering game. Across 24 matches pairing 23 expert humans with 16 AI agents, 387 delegation decisions and 1440 adoption decisions were recorded. The central claims are that human-AI teams outperform either humans or AI alone, yet humans exhibit suboptimal reliance patterns: under-reliance on correct AI suggestions in 3.9% of opportunities and over-reliance on incorrect AI suggestions in 1.7% of cases. Additional findings include model confidence near chance level on disagreements and confirmation bias producing 64.5% under-reliance when AI output matches an initial human error. The authors recommend calibrated confidence, evidence-grounded explanations, and trust-refinement mechanisms.","tokens_in":1907,"tokens_out":615,"duration_ms":30925,"significance":"If the reported reliance rates and bias patterns hold, the work supplies concrete empirical data on when and why humans delegate or adopt AI output in a high-stakes collaborative setting. The design that measures both delegation and adoption decisions from the same participants is a clear strength, as is the scale of observed decisions. The findings could inform interface design for better-calibrated human-AI teams. However, the competitive expert QA context limits immediate generalizability, and the absence of statistical tests on the key percentages weakens the evidential basis for labeling the behavior 'suboptimal.'","major_comments":[{"comment":"Abstract: the central claim that humans 'make suboptimal collaboration decisions' is supported only by the raw percentages 3.9% and 1.7%; no statistical tests, confidence intervals, or comparison to a normative baseline (e.g., expected error under random or optimal policy) are reported, leaving open whether these rates differ reliably from chance or from optimal play within the game.","section":"Abstract"},{"comment":"Abstract and Discussion: the inference that the observed under- and over-reliance rates demonstrate generally suboptimal human decision-making is load-bearing for the paper's contribution, yet the design is confined to a competitive scoring game with expert participants and 16 specific AI agents; no within-paper comparison or robustness check tests whether the same direction or magnitude of bias appears in non-competitive tasks, with non-experts, or on different question distributions.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract states that 'reported model confidence is near chance when humans and AI disagree' but does not indicate how confidence was elicited from the AI agents or whether it was normalized across the 16 models.","section":"Abstract"},{"comment":"Methods section (inferred from sample sizes): exclusion criteria, inter-rater reliability for decision coding, and any controls for order or fatigue effects across the 24 matches are not mentioned in the provided abstract, which would aid reproducibility.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the evidential basis for our claims and the scope of generalizability. We address each major comment below and outline planned revisions.","responses":[{"response":"We agree that statistical support would strengthen the central claim. In the revised manuscript we will add binomial confidence intervals around the reported 3.9% under-reliance and 1.7% over-reliance rates. We will also include explicit comparisons of these rates against (a) a random policy baseline derived from the observed human and AI accuracies and (b) an optimal policy that always follows the higher-accuracy agent on each question. These additions will clarify whether the observed deviations are statistically distinguishable from chance or optimal behavior within the game.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that humans 'make suboptimal collaboration decisions' is supported only by the raw percentages 3.9% and 1.7%; no statistical tests, confidence intervals, or comparison to a normative baseline (e.g., expected error under random or optimal policy) are reported, leaving open whether these rates differ reliably from chance or from optimal play within the game."},{"response":"We acknowledge that the study is limited to a competitive expert QA setting with the 16 AI agents used. The manuscript does not assert that the observed bias magnitudes are universal. In revision we will qualify the abstract and discussion to state explicitly that the findings are tied to this high-stakes collaborative game and to recommend future work examining non-competitive tasks, non-expert participants, and varied question distributions. Because the current work is an observational study of existing matches, we cannot add new within-paper robustness experiments without collecting additional data.","revision_made":"partial","referee_comment":"[Abstract] Abstract and Discussion: the inference that the observed under- and over-reliance rates demonstrate generally suboptimal human decision-making is load-bearing for the paper's contribution, yet the design is confined to a competitive scoring game with expert participants and 16 specific AI agents; no within-paper comparison or robustness check tests whether the same direction or magnitude of bias appears in non-competitive tasks, with non-experts, or on different question distributions."}],"tokens_in":1548,"tokens_out":476,"duration_ms":24045,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The useful part is that they tracked both delegation choices (handing off before seeing the answer) and adoption choices (taking or rejecting a shown suggestion) from the same 23 expert humans across 24 matches against 16 AI agents. That yields 387 delegation decisions and 1440 adoption ones, and the abstract states this combination is rare in prior work. They also report that the combined human-AI teams beat either alone, with under-reliance on correct AI at 3.9% and over-reliance on wrong AI at 1.7%, plus a 64.5% confirmation-bias rate when AI echoed an initial human error.\n\nThose numbers come directly from recorded decisions in the game, so there is no circularity from fitted models. The design lets them observe real trade-offs inside one cohort rather than stitching separate studies.\n\nThe soft spots sit in the interpretation and scope. The abstract supplies raw percentages but no statistical tests, exclusion rules, or checks for how the competitive scoring might push reliance one way or another. More importantly, the central claim that humans make generally suboptimal collaboration decisions rests on rates measured only with experts in this scoring game. Nothing in the reported design tests whether the same 3.9% and 1.7% figures, or even the same direction of bias, appear with non-experts or on non-competitive tasks. That gap matches the stress-test concern exactly.\n\nThe work is aimed at researchers who study human-AI reliance patterns and want empirical baselines from a single realistic session. It deserves a serious referee because the joint measurement is a concrete step forward even if the methods section needs more detail on controls and the discussion needs to qualify the scope of the suboptimal claim.","headline":"The paper gives joint counts on delegation and adoption from the same users in one game and reports low under- and over-reliance rates, but the competitive expert setup leaves generalizability open.","tokens_in":2408,"tokens_out":428,"would_cite":false,"duration_ms":23469,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Humans in a question-answering game with AI outperform either alone but still miss correct AI suggestions and accept wrong ones due to bias and uncalibrated confidence.","keywords":["human-AI collaboration","delegation","adoption","trust","question answering","confirmation bias","reliance decisions","over-reliance"],"falsifier":"Repeating the same measures of under-reliance and over-reliance in a non-competitive, open-ended question-answering task with non-expert users would directly test whether the reported rates and bias patterns hold outside the game setting.","tokens_in":2680,"feed_emoji":"🤖","tokens_out":813,"duration_ms":23724,"temperature":0.7,"pith_summary":"The paper studies how people decide to delegate tasks to AI without seeing its output and how they decide to adopt or reject AI suggestions after seeing them. These two reliance patterns are examined together in the same users during a competitive question-answering game that pairs expert humans with multiple AI agents. Collaboration improves overall accuracy compared with solo humans or solo AI, yet humans still miss 3.9 percent of opportunities to use correct AI answers and follow misleading AI answers 1.7 percent of the time. Confirmation bias raises under-reliance to 64.5 percent when an AI suggestion matches a human's initial wrong answer, and reported AI confidence is near chance level on disagreements. The findings point to concrete design changes such as better calibrated confidence scores and evidence-grounded explanations.","feed_headline":"Humans miss 3.9% of correct AI suggestions in QA game","feed_subtitle":"Teams beat solo humans or AI, yet confirmation bias and weak confidence signals still produce under-reliance and over-reliance.","key_machinery":"The separation of delegation choice (deciding to let AI act autonomously without seeing its output) from adoption choice (evaluating a visible AI suggestion), measured together for the same users inside a competitive question-answering game.","core_discovery":"In 24 matches that produced 387 delegation decisions and 1440 adoption decisions, human-AI teams performed better than either humans or AI working alone, yet humans under-relied on correct AI suggestions in 3.9 percent of opportunities and over-relied on misleading AI suggestions in 1.7 percent of cases. When humans and AI disagreed, reported model confidence performed near chance. Confirmation bias produced markedly higher under-reliance (64.5 percent) precisely when an AI suggestion agreed with a human's initial incorrect answer. Both parties therefore contributed errors, and the study recommends calibrated confidence, evidence-grounded explanations, and mechanisms that help users refine t","pith_inferences":["If the observed rates persist outside games, user training focused on detecting confirmation bias could measurably improve reliance accuracy.","Interfaces that surface disagreement more visibly than agreement might reduce the 64.5 percent under-reliance spike.","The same delegation-versus-adoption split could be measured in domains such as medical diagnosis or legal review to check whether the same bias patterns appear.","Allowing users to adjust an AI's displayed confidence threshold over repeated interactions offers a testable way to close the performance gap."],"forward_implications":["Human-AI teams reach higher accuracy than solo humans or solo AI in the same question-answering task.","Suboptimal reliance appears in both directions: missed correct AI suggestions and accepted incorrect ones.","Confirmation bias raises under-reliance sharply when AI output matches an initial human error.","AI confidence scores are near chance level precisely when humans and AI disagree.","Calibrated confidence, evidence-based explanations, and trust-refinement tools are proposed as direct remedies."],"fun_headline_variants":["3.9% of correct AI suggestions under-relied on in human QA game","64.5% under-reliance from confirmation bias when AI agrees with error","AI confidence near chance level amid human disagreements in matches","1.7% over-reliance on misleading AI suggestions during QA adoption"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The competitive game format with expert humans and the chosen AI agents produces reliance decisions that generalize to other human-AI collaboration settings.","fun_headline_variants_meta":{"raw":{"variants":["3.9% of correct AI suggestions under-relied on in human QA game","64.5% under-reliance from confirmation bias when AI agrees with error","AI confidence near chance level amid human disagreements in matches","1.7% over-reliance on misleading AI suggestions during QA adoption"]},"model":"grok-4.3","cost_usd":0.007472,"raw_usage":{"total_tokens":3482,"prompt_tokens":771,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":74724500,"prompt_tokens_details":{"text_tokens":771,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2637,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":771,"tokens_out":74,"duration_ms":21062,"temperature":1.0,"reasoning_tokens":2637,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T12:11:55.058822+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the same measures of under-reliance and over-reliance in a non-competitive, open-ended question-answering task with non-expert users would directly test whether the reported rates and bias patterns hold outside the game setting.","supporting_citations":[],"review_version":1}