{"id":"63910e03-d6fc-424f-9f65-2be1c1c8a3ce","arxiv_id":"2411.08583","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A pre-registered experiment found that an AI providing only pro and con evidence, without recommendations, did not improve decision performance and was used shallowly by participants.","lead":"This experiment tested an AI assistant that offers pros and cons instead of direct recommendations, and found it did not improve people's decision accuracy. It is a useful counterpoint to optimistic claims that 'hypothesis-driven' AI will automatically make human-AI collaboration better.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Evaluative AI null result may reflect an unvalidated, possibly unusable evidence display: SHAP log-odds contributions are not mapped to the probability scale participants must estimate, so the test may not instantiate Miller's framework.","rationale":"The reader's weakest assumption correctly identifies the evidence implementation as the key vulnerability. I sharpen it: beyond accuracy of the GPT-4o text, the evidence format may be fundamentally insufficient for the specific probability-estimation task, because SHAP contributions are on a log-odds scale and are not accompanied by the model's intercept, base rate, or aggregation rule. In Recommendation Only, the AI's output probability is shown directly and that group is the only one to beat the control; in Evaluative AI and Evidence Only the probability is hidden, so any deficit could be due to missing calibration information rather than to the hypothesis-driven evidence concept. This concern is acknowledged in the paper's limitations, which say the evidence should be clearer and more accessible. It does not invalidate the internally valid null result for this implementation, but it does weaken the broader claim that the Evaluative AI framework did not improve decision-making. Because the authors already report the result as specific to their experiment and propose future work on evidence presentation, the reader's CONDITIONAL verdict is appropriate; no verdict change is needed, though the revision should either add evidence-comprehension validation or explicitly narrow the claim to the tested implementation.","tokens_in":18488,"tokens_out":9411,"duration_ms":91466,"concrete_test":"Run a focused manipulation check using the exact evidence screens from the four experimental individuals: show them to a fresh sample of participants and ask (1) 'According to the AI, which features are evidence for above-median income?' and (2) 'What probability do you think the AI would assign to this person?' If a substantial fraction cannot identify the intended probability or correctly aggregate the pro and con evidence, the display is not a fair test of the framework. Alternatively, add an Evaluative AI condition with an explicit calibration cue (e.g., 'Starting from 50%, these features shift the estimate up or down') and compare performance to the original Evaluative AI group; if the cue improves Brier scores, the original null is an artifact of a missing mapping from evidence to probability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the assumption that the SHAP/GPT-4o pro-con evidence shown in the Evidence Only and Evaluative AI conditions is a valid operationalization of Miller's evaluative evidence, and that participants can use it to make the probability judgments the task requires. This assumption is not established. Section 3.7 describes converting SHAP values to text via GPT-4o with no reported validation of accuracy or comprehensibility, and the comprehension checks in Section 3.5 cover instructions, not the evidence display. The pretest noted in Section 1 found bar charts were poorly understood, so text was added, but no check confirms the text actually fixes the problem. More fundamentally, the dependent variable is a calibrated probability measured by Brier score, while SHAP values are feature-level log-odds contributions with no stated baseline or aggregation rule. Without the model's probability or an explicit base-rate cue, the evidence tells participants which features point up or down but not how to map those contributions to a percentage estimate. Recommendation Only, which provides the AI's probability, is the only group that beat the humans-only baseline, while Evaluative AI, which hides that probability, performs worst numerically. The paper's own Limitations section concedes that pro-con evidence presentation needs improvement. Thus the null result may describe this particular implementation rather than the Evaluative AI framework, and the load-bearing assumption of implementation fidelity is the weakest point in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a preregistered, incentivized online experiment (N = 375) testing Miller's Evaluative AI framework. Participants estimated the probability that each of four individuals had a net income above the median, with five between-subjects conditions: no AI, AI recommendation only, evidence only, recommendation plus evidence, and the Evaluative AI condition in which pro/con evidence was available on demand. The central findings are that Brier-score performance did not differ significantly across conditions, with the Evaluative AI condition numerically worst; evidence-presenting conditions were slower than control or recommendation-only; NASA-TLX cognitive load did not differ; and qualitative coding suggested that AI-assisted participants mentioned fewer features. The authors conclude that the Evaluative AI framework, at least as implemented here, did not improve decision-making performance and that engagement with evidence was limited.","tokens_in":18761,"tokens_out":5310,"duration_ms":47043,"significance":"If the null result is taken at face value, it is a useful empirical counterpoint to Le et al. (2024a) and to Miller's theoretical claims, and it provides a cautionary data point for hypothesis-driven XAI. The study has real methodological strengths: it is preregistered, uses a power analysis with Bonferroni correction, applies non-parametric tests with a control group, uses monetary incentives tied to Brier score, and makes data and code available. The main weakness is that the validity of the evidence operationalization is not established, which makes the scope of the conclusion uncertain. The study is therefore worth publishing only after the authors either validate the evidence display or explicitly restrict their claims to the specific implementation tested.","major_comments":[{"comment":"Section 3.7 and Section 3.2: The pro/con evidence is generated from SHAP values and converted to free text via GPT-4o, but the manuscript provides no validation that these statements accurately reflect the SHAP contributions, no comprehensibility or manipulation check, and no evidence that participants could map the directional statements to the probability scale they had to report. The pretest mentioned in Section 1 found bar charts poorly understood, yet no check confirms that the added textual descriptions fixed the problem. Because the central null result in the Evidence Only and Evaluative AI conditions depends on participants actually receiving and understanding Miller-style evaluative evidence, the study currently cannot distinguish 'the framework does not improve decisions' from 'this operationalization of evidence was not usable.' I recommend either a comprehension check on the evidence display, an audit of the GPT-4o-generated statements, or a condition in which the evidence is coupled with the AI's baseline probability.","section":"Section 3.7"},{"comment":"Section 3.3 (Brier score) and Section 3.7: The outcome is a calibrated probability, but the evaluative evidence is presented as feature-level SHAP contributions without any stated mapping from those log-odds-style contributions to the percentage estimates participants must produce. Evidence Only and Evaluative AI participants never see the AI's probability estimate or a base-rate anchor, whereas Recommendation Only participants do. This confounds the presence of a recommendation with the availability of a response-scale cue; the observed advantage of Recommendation Only and the poor numerical performance of Evaluative AI may reflect the absence of a probability anchor rather than the absence of a recommendation. The paper should either provide the model's predicted probability or an explicit aggregation instruction in the evidence conditions, or restrict the conclusion to the specific evidence presentation tested.","section":"Section 3.3"},{"comment":"Section 3.4 (Decision-Making Process) and Section 4.4: The qualitative coding is described as performed by the author with the assistance of GPT-4o, but no inter-rater reliability, double-coding subsample, or validation of the LLM-assisted coding is reported. The claims about AI mentions, feature mentions, and the resulting cognitive-offloading interpretation rest on these categories. A second human coder on a random subsample with a kappa statistic, or a documented validation of the GPT-4o coding protocol, is necessary to support those conclusions.","section":"Section 3.4"},{"comment":"The abstract characterizes engagement as 'limited,' but Section 4.5 reports that 62.64% of Evaluative AI participants clicked on both the pro and con evidence in all four rounds (i.e., all eight evidence items). If the intended claim is that participants did not deeply process the evidence, the click data alone do not support that; the paper should state the specific measure on which 'limited engagement' is based and reconcile the high click-through rate with the superficial-engagement interpretation.","section":"Abstract and Section 4.5"}],"minor_comments":[{"comment":"There is a typo in the subsection heading ('Desicion-Making Process'); similarly, 'Apendix' appears in Section 3.2.","section":"Section 3.5"},{"comment":"Reporting effect sizes and confidence intervals for the pairwise comparisons (or at least the test statistics) would improve interpretability, since Kruskal-Wallis p-values alone do not convey the magnitude of the null result.","section":"Section 4.1"},{"comment":"Appendix A states the AI is correct 77% of the time, whereas Section 3.7 reports 15/20 (75%) accuracy with a 50% threshold; these numbers should be reconciled.","section":"Appendix A"},{"comment":"The final group sizes are noticeably uneven (62 to 91 participants); the manuscript should state how randomization was implemented and whether the imbalance was checked.","section":"Section 3.8"},{"comment":"The pairwise chi-squared tests are reported only through p-values; the test statistics or exact p-values with the correction method should be given.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Tim — quick take on 2411.08583. It's the second empirical test of Miller's Evaluative AI framework and the first to report a null result. Pre-registered, powered to n=375, five between-subject conditions including a no-AI control, income-estimation task with Brier-score incentives, click tracking for the opt-in evidence. Recommendation Only beat the no-AI baseline; the Evaluative AI group performed worst numerically, and overall there was no significant performance difference. That contrast with Le et al. (2024a) is the interesting finding.\n\nThe paper is honest and methodical. Hypotheses are preregistered, analysis is sensible (Kruskal-Wallis with Bonferroni, non-normal data), and the qualitative analysis gives some color, though the GPT-4o-assisted coding has no reported inter-rater reliability. Data and code are promised in a public repository; the manuscript itself doesn't give a direct link, which should be fixed.\n\nThe soft spot is the one the stress test flags: implementation fidelity. The dependent variable is a calibrated probability, but the evidence shown to participants is SHAP log-odds contributions, never mapped to a probability or given a base-rate anchor. Without that mapping, \"this feature points up\" doesn't tell someone how to move their 50% prior. The GPT-4o text was never validated for comprehensibility, and the comprehension checks covered instructions, not the evidence display. The paper's own Limitations section concedes the evidence presentation needs work. So the null result is fairly described as \"this implementation of the framework did not help,\" not as a decisive falsification of Miller's idea. The author actually says that, more or less, which I respect.\n\nThere is also the point that the AI model itself was only about 75% accurate, roughly at human level; 43% of control participants beat the AI. So low headroom may explain the null as much as the framework does. The author discusses this, so it's not hidden.\n\nI think the reader's conditional verdict is right. The weaknesses are addressable and don't undermine the preregistered design. The paper deserves full peer review. For a reading group it would be a good discussion piece precisely because the operationalization question matters for the whole XAI evaluation literature.\n\nBottom line: send it to reviewers. It's a legitimate, reproducible empirical contribution even if the interpretation is bounded.","headline":"A pre-registered, powered null result for Miller's Evaluative AI framework that is honest about its own limitations; the main caveat is whether the SHAP-based evidence display actually instantiates the framework.","tokens_in":19269,"tokens_out":2009,"would_cite":true,"duration_ms":18634,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a 375-participant experiment, replacing AI recommendations with optional pro and con evidence did not improve decision quality or engagement, and users offloaded cognition much as they do with traditional AI.","keywords":["Evaluative AI","explainable AI","human-AI decision-making","pro and con evidence","Brier score","cognitive offloading","behavioral experiment","SHAP"],"falsifier":"Re-run the same five-condition experiment with an evidence-comprehension check that participants must pass before deciding, and with the Evaluative AI group's bonus tied to using the evidence; if that group's Brier score drops significantly below the control group's, the claim that this framework does not improve decisions would be overturned in that setting.","tokens_in":18310,"feed_emoji":"🤖","tokens_out":9397,"duration_ms":79730,"temperature":0.7,"pith_summary":"This paper sets out to test the Evaluative AI framework's core promise: that presenting decision-makers with pro and con evidence for each option, rather than an AI recommendation, will lead to more informed decisions and less blind reliance. The experiment used an incentivized income-estimation task in which 375 lay participants judged whether four people earned above-median net income, across five conditions: no assistance, recommendation only, evidence only, recommendation plus evidence, and the framework's optional-evidence design. The result was a null effect: the Evaluative AI condition's Brier score did not differ significantly from any other condition and was numerically the worst of the five. Evidence-displaying conditions were significantly slower, cognitive load did not rise, and qualitative reports indicated that AI-assisted participants relied on fewer features than unsupported participants, which the paper reads as cognitive offloading and potential automation bias rather than deeper engagement. If this result holds, it undercuts the empirical case that a recommendation-free, hypothesis-driven interface alone can improve human-AI decision making, and it shifts attention to the quality and usability of the evidence itself.","feed_headline":"AI that argues both sides did not improve decisions","feed_subtitle":"In an income-estimation experiment, the Evaluative AI group scored no better than people working alone.","key_machinery":"The central machinery is a five-condition randomized experiment built on the contrast between recommendation-driven and hypothesis-driven support. In the Evaluative AI condition, participants see only two buttons, one for pro evidence and one for con evidence, and can open them at any time; no recommendation is ever offered. The evidence itself is generated from SHAP feature attributions of a logistic regression model (each personal characteristic contributes a positive or negative push to the predicted probability), displayed as bar charts and converted into short textual arguments. Performance is scored with the Brier score, $\\frac{1}{N}\\sum_{i=1}^{N}(p_i-o_i)^2$, which also determines the bonus payment; decision time, NASA-TLX cognitive load, and qualitative coding of free-text decision descriptions complete the measurement of how users actually reasoned.","core_discovery":"The central discovery is a null result, stated on the paper's own terms: in an incentivized between-subjects experiment with 375 lay participants estimating whether each of four individuals earns above-median net income, the treatment that implemented the Evaluative AI framework—pro and con evidence available on request, with no recommendation—did not improve decision-making performance. The mean Brier score $\\frac{1}{N}\\sum_{i=1}^{N}(p_i-o_i)^2$ for the Evaluative AI group was $0.230$, the worst among the five groups, and a Kruskal-Wallis test found no significant between-group difference ($p=0.154$). The paper also reports limited engagement: 62.64% of Evaluative AI participants clicked both evidence buttons in every round, and qualitative descriptions of the decision process showed that AI-supported participants relied on fewer features than unsupported controls, which the author interprets as cognitive offloading and potential automation bias rather than deeper reasoning with the evidence. The conclusion is that this implementation of the framework did not deliver the performance or engagement gains it was designed to produce.","pith_inferences":["The null result may be an implementation ceiling: if the textual pro/con evidence was hard to parse or did not match participants' mental models, the framework itself remains untested. A comprehension check on the evidence would separate these possibilities.","The incentive design may matter: a £3 show-up fee with a bonus up to £6 per four tasks may not create the high-stakes, low-frequency setting the framework targets. A field-like task with real professional consequences could produce different engagement.","This result fits a cost-benefit view of overreliance: when optional evidence costs extra clicks and reading time, users may rationally skip it. A design that makes evidence an unavoidable first step, or that gives users a personal stake in accuracy, could recover the intended effect.","The binary income task may be too simple for pros-and-cons reasoning to add value; the framework's natural habitat is multi-hypothesis diagnosis where several options compete, so a multi-class extension is a direct next test."],"forward_implications":["No AI support condition significantly beat doing the task alone on Brier score; the Evaluative AI condition's mean score of $0.230$ was numerically the worst, so the framework's hoped-for performance gain did not materialize.","Evidence on demand slowed decisions: Evaluative AI (mean 51.7s), Evidence Only (56.4s), and Recommendation and Evidence (57.2s) all took significantly longer than Control or Recommendation Only (about 41s).","Cognitive load, as measured by NASA-TLX, did not differ between any treatments, contradicting the hypothesis that hypothesis-driven support would demand more mental effort.","Qualitative reports indicate all AI-assisted groups mentioned individual features less often than the no-AI control, which the paper reads as cognitive offloading and potential automation bias rather than deeper engagement.","Only 62.64% of participants in the Evaluative AI condition used the optional evidence in every round, so superficial engagement persisted even when evidence was the only AI feature."],"supporting_citations":[{"why":"Defines the Evaluative AI framework and the hypothesis-driven, pro/con evidence approach that the experiment is designed to test.","marker":"Miller (2023)"},{"why":"The prior empirical study that reported improved decision quality from a hypothesis-driven approach; this paper's null result is the direct contrast.","marker":"Le et al. (2024a)"},{"why":"Provides the SHAP method used to derive pro and con evidence from the logistic regression model.","marker":"Lundberg and Lee (2017)"},{"why":"Documents the engagement problem with explanations and cognitive forcing functions, used to interpret limited evidence use and overreliance.","marker":"Bucinca et al. (2021)"},{"why":"Supplies the cost-benefit account of overreliance that the paper uses to explain why participants might skip optional evidence.","marker":"Vasconcelos et al. (2023)"},{"why":"Prior finding that recommendations help performance but explanations add little, used to contextualize the absent improvement.","marker":"Bansal et al. (2021)"},{"why":"Earlier evidence that explanations alone can slightly improve user performance, a benchmark this study's null result does not replicate.","marker":"Lai and Tan (2019)"},{"why":"Shows that even data scientists struggle with SHAP-style bar charts, motivating the textual evidence and flagging a possible implementation problem.","marker":"Kaur et al. (2020)"}],"fun_headline_variants":["Evaluative AI fails to improve decisions in test","Pro-con AI evidence no better than no AI","Both-sides AI: no decision gains, limited engagement","AI arguing both sides fails to help decisions","Evaluative AI null result on decision quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pro and con evidence participants could click on was clear enough and faithful enough to what the framework intends that the experiment actually tested the framework's idea rather than a flawed presentation of it.","fun_headline_variants_meta":{"raw":{"variants":["Evaluative AI fails to improve decisions in test","Pro-con AI evidence no better than no AI","Both-sides AI: no decision gains, limited engagement","AI arguing both sides fails to help decisions","Evaluative AI null result on decision quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00028,"raw_usage":{"total_tokens":1608,"prompt_tokens":842,"completion_tokens":766,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":695}},"tokens_in":458,"tokens_out":766,"duration_ms":8015,"temperature":1.0,"reasoning_tokens":695,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:30:51.923476+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same five-condition experiment with an evidence-comprehension check that participants must pass before deciding, and with the Evaluative AI group's bonus tied to using the evidence; if that group's Brier score drops significantly below the control group's, the claim that this framework does not improve decisions would be overturned in that setting.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Evaluative AI framework and the hypothesis-driven, pro/con evidence approach that the experiment is designed to test."},{"cited_title":"S., and Krishna, R","cited_arxiv_id":null,"evidence_quote":"Supplies the cost-benefit account of overreliance that the paper uses to explain why participants might skip optional evidence."},{"cited_title":"and Tan, C","cited_arxiv_id":null,"evidence_quote":"Earlier evidence that explanations alone can slightly improve user performance, a benchmark this study's null result does not replicate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that even data scientists struggle with SHAP-style bar charts, motivating the textual evidence and flagging a possible implementation problem."}],"review_version":1}