{"id":"92982bde-69e0-47f2-be7f-e0d29d9e6d74","arxiv_id":"2607.05680","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI attribution has no net effect on perceived reasonableness of controversial legal advice because objectivity gains are offset by losses in comprehensiveness and attention to special circumstances; reasoning helps via objectivity.","lead":"A preregistered experiment with 3,348 Chinese adults finds no net difference in how reasonable people rate identical, legally correct but socially controversial legal advice when it is labeled AI versus human. Opposing perceptions of AI objectivity versus missing context cancel out, while adding legal reasoning raises acceptance for both sources.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Mediation paths that explain the null are not causally identified; single-item post-treatment mediators leave the paper's mechanistic claim under-supported.","rationale":"The reader's weakest_assumption correctly isolates the load-bearing soft spot: the null main effect is well-supported by the preregistered factorial design and large N, but the paper's theoretical payoff—the cancelling mediation story—depends on unmanipulated single-item mediators and the standard no-confounding assumption the authors themselves flag. My concern is essentially the same, sharpened only by noting that the abstract and discussion present the opposing paths as the explanation of the null, so identification failure would leave the central narrative under-supported even while the ATE claim stands. No stronger internal inconsistency appears; cultural/vignette limits and missing public data are secondary. Verdict therefore stays CONDITIONAL, pending either public data that allow the residualization check above or multi-item, preferably pre-treatment, validation of the mediators.","tokens_in":16756,"tokens_out":548,"duration_ms":5387,"concrete_test":"Re-estimate the attribution mediation model after residualizing each of the four mediators on a pre-treatment measure of general AI stereotypes (or, if unavailable, on the collected 'perceptions of the reliability of AI legal software' item) and on the other three mediators; if the indirect effects through objectivity, comprehensiveness, and special-circumstances shrink below significance or change sign, the claimed opposing pathways are not identified as advice-specific processes.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's distinctive contribution is not the null main effect of AI attribution (which is cleanly identified by the 2\times2 design) but the claim that this null is produced by opposing psychological pathways: AI raises perceived objectivity (positive) while lowering comprehensiveness and attention to special circumstances (negative). That claim rests on simultaneous SEM mediation of four single-item, post-treatment measures (§5.2.5, Tables 3–4). The authors correctly note that causal interpretation requires no unmeasured confounding of mediator–outcome links and that the items may partly reflect source stereotypes rather than advice-specific evaluations. Because the mediators were never experimentally manipulated, and because they are single items collected after the source label, the opposing-path story is correlational. If the mediators primarily capture pre-existing AI stereotypes activated by the label, the paper has documented stereotype activation rather than the process by which people evaluate the advice itself. The null ATE remains intact; the mechanistic explanation that the abstract and discussion foreground does not.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports a preregistered 2×2 survey experiment (N=3,348 after comprehension checks) in mainland China testing how laypeople rate the reasonableness of identical, legally correct but socially controversial legal advice when attributed to an AI system versus a human lawyer, and when accompanied by legal reasoning or not. Three pretested civil scenarios (betrothal gift, estate division, pre-existing medical condition) are used. Attribution has no net effect on reasonableness (pooled ATE ≈ 0.012, p=0.801). Simultaneous SEM mediation with four single-item post-treatment mediators indicates opposing paths: AI attribution raises perceived objectivity (positive indirect effect) while lowering comprehensiveness and attention to special circumstances (negative indirect effects). Providing reasoning raises reasonableness regardless of source, largely via objectivity. Qualitative justifications are presented as corroboration. The authors conclude that responses to AI legal advisors reflect balancing of competing normative expectations rather than rigid algorithm aversion.","tokens_in":16972,"tokens_out":1186,"duration_ms":9633,"significance":"If the main experimental results hold, the paper makes a useful contribution to algorithm-aversion research in a high-stakes, normatively charged domain. The clean identification of a null attribution effect for legally correct but socially controversial advice, the large preregistered sample, the use of pretested scenarios, HC2 robust SEs, and the robust positive effect of reasoning are genuine strengths. The design isolates source label while holding advice content fixed, which prior legal-advice studies often do not. The qualitative material adds interpretive texture. Even if the mediation paths remain only partially identified, documenting that a null ATE can mask opposing evaluative dimensions is informative for theory and for the design of AI legal interfaces. The China setting also expands geographic coverage of a literature that is still heavily Western.","major_comments":[{"comment":"§5.2.5 and Tables 3–4: The paper’s distinctive mechanistic claim—that the null attribution effect is produced by opposing pathways through objectivity versus comprehensiveness/attention to special circumstances—rests on simultaneous SEM mediation of four single-item, post-treatment mediators. The authors correctly note the no-unmeasured-confounding assumption and that the items may partly reflect source stereotypes activated by the label rather than advice-specific evaluations. Because the mediators were never manipulated, the opposing-path story is correlational. The abstract, introduction, and discussion currently present these paths as the explanation of the null. Either (a) reframe the mediation as exploratory/descriptive evidence of associated perceptions, or (b) add sensitivity analyses (e.g., VanderWeele-style bounds) and substantially qualify causal language so that the identifie","section":null},{"comment":"§5.1.3 and §4: The three scenarios were selected for approximate 60–40 opinion splits, yet the reasoning effect is heterogeneous (ATE ≈ 0.35 and 0.52 in two scenarios, near zero in the pre-existing medical condition scenario; §5.2.2). The paper pools for the main claims and does not systematically analyze why reasoning fails in one case. Given that scenario selection is an explicit design choice, the manuscript should either (a) report and discuss scenario-level heterogeneity as a substantive finding about when explanations help, or (b) justify pooling more carefully and treat the null reasoning effect in one scenario as a boundary condition rather than noise.","section":null}],"minor_comments":[{"comment":"Figure 2 caption refers to “three case scenarios” while the surrounding text and figure content cover six cases from two surveys; align caption and text.","section":null},{"comment":"Table 1 reports pairwise correlations and descriptives; consider adding a brief note on whether mediator intercorrelations (e.g., comprehensiveness and attention to special circumstances r=0.52) raise concerns about discriminant validity of the single items.","section":null},{"comment":"§5.2.5 footnote: the switch from the preregistered Imai–Keele–Yamamoto framework to simultaneous SEM is disclosed and results are said to be qualitatively similar; a short appendix table comparing the two would strengthen transparency.","section":null},{"comment":"Qualitative excerpts in §6 are illustrative but unsystematic; a brief coding scheme or frequency counts for the objectivity vs. contextual-sensitivity themes would make the corroboration claim more transparent.","section":null},{"comment":"Minor typos and wording: “reasona bleness” (abstract), “He’s signature” vs. Yu in the traffic vignette description (§4.1.1), and occasional tense/number inconsistencies in the mediation write-up.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The experimental core (null attribution ATE, positive reasoning ATE, preregistration, large N) is solid and publishable with modest revision. The main risk is over-claiming on the mediation story; if the authors reframe that material as exploratory, the paper is close to ready. Fit for a general HCI / computational social science / law-and-technology venue is good; pure methods journals might want stronger identification of the paths."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The headline result is solid: when the legal advice text is held fixed and is legally correct but socially contested, Chinese lay participants rate AI-attributed and lawyer-attributed advice as equally reasonable (pooled ATE ~0). Reasoning helps a lot, mainly by raising perceived objectivity. That is the clean, preregistered contribution.\n\nWhat is new is the design package. They pretested cases for controversy, held the advice constant across source labels, ran a large N after comprehension checks (3,348), used HC2 SEs, and reported both main effects and simultaneous SEM mediation with covariates. The qualitative open responses line up with the quantitative tension between objectivity and contextual sensitivity. Citation pattern is appropriate; self-cites are limited and not load-bearing. Math and estimation look standard and transparent.\n\nThe soft spot is exactly where the stress-test puts it, and the authors already flag it. The paper’s distinctive claim is not the null ATE but the canceling pathways (objectivity up, comprehensiveness and special-circumstances attention down). Those rest on four single-item, post-treatment mediators that were never manipulated. Causal mediation is not identified; the paths could partly be stereotype activation by the source label rather than evaluation of the advice itself. That weakens the mechanistic story the abstract and discussion foreground, but it does not touch the null main effect or the reasoning effect. Single-item measures and China-only civil vignettes are real limits; they are also clearly stated.\n\nThis is for people working on algorithm aversion, legal-tech HCI, and explanation design in normatively loaded domains. It is not a paradigm shift, but it is careful empirical work that advances the literature without overclaiming. I would send it to peer review. Data/code are promised on OSF; once public, the main estimates should be easy to check. Worth engaging.","headline":"Clean null on AI vs human source for fixed-text controversial legal advice; the opposing-mediation story is interesting but only correlational.","tokens_in":17549,"tokens_out":469,"would_cite":true,"duration_ms":4868,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"People rate AI and human legal advice as equally reasonable; opposing judgments of objectivity and context cancel out.","keywords":["algorithmic aversion","legal advice","attribution","reasoning","objectivity","uniqueness neglect","mediation"],"falsifier":"Re-run the same vignettes with multi-item validated scales for objectivity, comprehensiveness, and uniqueness neglect, or experimentally manipulate those mediators directly; if the opposing indirect effects disappear while the null main effect of source remains, the paper’s mechanistic account is wrong.","tokens_in":17619,"feed_emoji":"⚖️","tokens_out":562,"duration_ms":5309,"temperature":0.7,"pith_summary":"This preregistered experiment with 3,348 adults in mainland China asks whether laypeople accept legally correct but socially controversial advice when it is labelled as coming from an AI system rather than a human lawyer. Attribution has no net effect on how reasonable the advice seems. That null result is produced by cancelling pathways: AI-attributed advice is seen as more objective (which raises reasonableness) but less comprehensive and less attentive to special circumstances (which lowers it). Providing legal reasoning raises reasonableness for both sources, largely by boosting perceived objectivity. The paper argues that responses to machine legal advisors are not fixed algorithm aversion, but a balancing of competing expectations about neutrality and contextual sensitivity—and that this balance matters for how automated legal tools should be designed.","feed_headline":"AI and human legal advice rate equally reasonable","feed_subtitle":"Objectivity boosts acceptance; missing context cuts it; reasoning helps both.","key_machinery":"A 2×2 preregistered factorial survey experiment (source: AI vs human; reasoning present vs absent) on three pre-selected controversial Chinese civil-law scenarios, with simultaneous multi-mediator path models of four post-treatment perceptions (mistake, objectivity, comprehensiveness, attention to special circumstances).","core_discovery":"When identical, legally correct but socially controversial legal advice is attributed to an AI system rather than a human lawyer, average perceived reasonableness does not change. Mediation analysis shows the null is the product of opposing paths: higher perceived objectivity increases reasonableness, while lower perceived comprehensiveness and attention to special circumstances decrease it. Providing legal reasoning increases reasonableness for either source, mainly by raising perceived objectivity.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["AI legal advice rates equal to human on reasonableness","Opposing paths leave AI and human advice equally reasonable","Objectivity gains cancel context losses for AI legal advice","Reasoning boosts reasonableness of AI or human legal advice","AI and human legal advice: same scores, different perceptions"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claimed opposing pathways only hold if the four single-item ratings truly capture the intended perceptions and nothing unmeasured confounds how those perceptions relate to reasonableness.","fun_headline_variants_meta":{"raw":{"variants":["AI legal advice rates equal to human on reasonableness","Opposing paths leave AI and human advice equally reasonable","Objectivity gains cancel context losses for AI legal advice","Reasoning boosts reasonableness of AI or human legal advice","AI and human legal advice: same scores, different perceptions"]},"model":"grok-4.5","effort":"low","cost_usd":0.008228,"raw_usage":{"total_tokens":1914,"prompt_tokens":766,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":82280000,"prompt_tokens_details":{"text_tokens":766,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1082,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":766,"tokens_out":66,"duration_ms":8252,"temperature":1.0,"reasoning_tokens":1082,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T03:52:10.314921+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same vignettes with multi-item validated scales for objectivity, comprehensiveness, and uniqueness neglect, or experimentally manipulate those mediators directly; if the opposing indirect effects disappear while the null main effect of source remains, the paper’s mechanistic account is wrong.","supporting_citations":[],"review_version":1}