{"id":"10d91949-591a-4970-95fb-20a896b0961a","arxiv_id":"2406.05408","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"People project human-like difficulty judgments onto AI, causing overestimation on easy tasks, underestimation on hard ones, and distorted adoption even when AI excels selectively.","lead":"The paper formalizes and tests 'Human Projection,' where people judge AI using human difficulty and mistake-reasonableness frameworks, leading to biased performance estimates and all-or-nothing adoption. Smart generalists should read it because these biases can distort real-world decisions about deploying AI tools in domains like advice-giving.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Experimental isolation of projection mechanism from priors/exposure remains unverified as load-bearing assumption","rationale":"Reader's weakest_assumption matches the precise point where the causal attribution could fail; no other internal inconsistency (e.g., in formalization or field design) is detectable from provided material. Full-text review would be needed to confirm absence of controls, but the concern is load-bearing for the equilibrium-adoption predictions.","tokens_in":1670,"tokens_out":264,"duration_ms":9424,"concrete_test":"Re-analyze (or re-run) the main experiments with an added pre-survey measuring prior AI exposure and task priors, then test whether HP coefficients remain significant after including these as covariates; if effects attenuate below significance threshold, attribution to projection weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Central claim requires that belief distortions (overestimation on easy tasks, underestimation on hard, all-or-nothing adoption) are attributable to HP rather than task-specific priors or differential AI exposure. Abstract describes experiments and parenting-chatbot field setting but provides no detail on pre-treatment measurement or regression controls for these confounders; if participants enter with heterogeneous beliefs about AI jaggedness or task difficulty, observed patterns could arise without invoking projection of human difficulty frameworks.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"This paper studies Human Projection (HP): the tendency to evaluate AI using human frameworks for task difficulty and mistake reasonableness. It formalizes HP and its consequences for equilibrium adoption, then tests predictions experimentally. Key findings include projection of human difficulty onto AI (overestimating performance on easy tasks, underestimating on hard ones, and over-updating after easy failures/hard successes), all-or-nothing adoption even when AI outperforms humans on only some tasks, and a field experiment with a parenting-advice chatbot showing that less reasonable mistakes reduce trust and engagement more. Anthropomorphic design is argued to amplify these effects.","tokens_in":1756,"tokens_out":531,"duration_ms":19784,"significance":"If the results hold after addressing design details, the work offers a useful framework for understanding systematic biases in beliefs about AI capabilities and their impact on adoption. Strengths include the formalization of HP, the combination of lab experiments with a field test in a realistic setting, and the focus on design implications. These elements provide a foundation for further research on human-AI interaction.","major_comments":[{"comment":"Abstract: The description of multiple experiments and the parenting-chatbot field test provides no information on sample sizes, statistical power, pre-registration, or controls for alternative explanations such as prior AI exposure or task-specific priors. This is load-bearing for the central claim that belief distortions (overestimation on easy tasks, underestimation on hard ones, all-or-nothing adoption) are attributable to the HP mechanism rather than confounders.","section":"Abstract"},{"comment":"Experimental design sections: The assumption that the tasks and field setting isolate projection from other factors (e.g., heterogeneous beliefs about AI jaggedness or differential exposure) is unverified as load-bearing. Without pre-treatment measurement of priors or regression controls, observed patterns could arise independently of HP, weakening attribution to the proposed mechanism.","section":"Experimental design sections"}],"minor_comments":[{"comment":"Abstract: The summary of results is information-dense; consider separating the theoretical predictions from the empirical findings for improved readability.","section":"Abstract"},{"comment":"Notation and terminology: Ensure consistent use of 'HP' and related terms across the formalization and empirical sections to avoid ambiguity.","section":null}],"recommendation":"major_revision","confidential_remarks":"The lack of basic methodological details even in the abstract suggests the manuscript may be at an early stage; the journal's empirical standards will require substantial elaboration on design and robustness checks."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their detailed and constructive report. We address each major comment below. We agree that greater transparency on experimental details and robustness to alternative mechanisms will strengthen the manuscript and plan revisions accordingly.","responses":[{"response":"We agree that the abstract should convey more information on these elements to support attribution to the HP mechanism. The main text and appendix report sample sizes, power analyses, and pre-registration status for each study, along with discussion of design features intended to limit confounding. In revision we will expand the abstract to include summary sample sizes, note pre-registration, and reference the controls for prior exposure and task priors that are detailed in the experimental sections.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The description of multiple experiments and the parenting-chatbot field test provides no information on sample sizes, statistical power, pre-registration, or controls for alternative explanations such as prior AI exposure or task-specific priors. This is load-bearing for the central claim that belief distortions (overestimation on easy tasks, underestimation on hard ones, all-or-nothing adoption) are attributable to the HP mechanism rather than confounders."},{"response":"We acknowledge that explicit pre-treatment measurement of priors and additional regression controls would provide stronger evidence that the observed patterns are driven by HP rather than heterogeneous beliefs about AI capabilities or differential exposure. The current designs rely on randomization across conditions and careful task selection to isolate the mechanism, but we will revise the experimental sections to include pre-treatment elicitation of relevant priors and report regressions that control for prior AI exposure and task-specific beliefs. These additions will directly test robustness to the alternative explanations raised.","revision_made":"yes","referee_comment":"[Experimental design sections] Experimental design sections: The assumption that the tasks and field setting isolate projection from other factors (e.g., heterogeneous beliefs about AI jaggedness or differential exposure) is unverified as load-bearing. Without pre-treatment measurement of priors or regression controls, observed patterns could arise independently of HP, weakening attribution to the proposed mechanism."}],"tokens_in":1350,"tokens_out":443,"duration_ms":23911,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that people judge AI by projecting human task difficulty, overestimating on easy tasks and underestimating on hard ones. This produces systematic errors when AI performance is jagged and pushes adoption toward all-or-nothing patterns even when the AI is only better on some tasks. The field experiment with the parenting chatbot adds a real-world channel through engagement drops after unreasonable mistakes.","headline":"The paper formalizes human projection of difficulty onto AI, with lab and field tests showing distorted beliefs and all-or-nothing adoption, but the isolation from priors and exposure is the main unverified spot.","tokens_in":2227,"tokens_out":160,"would_cite":false,"duration_ms":13694,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Paper studies human projection bias in AI evaluation; RS derives physical constants from distinction logic — no shared machinery","alignment":"orthogonal","rationale":"The paper's core is an Ability Model (MLRP on ability/difficulty) plus feature projection (λ) leading to Berk-Nash equilibria and adoption distortions. This is standard behavioral-economics misspecification analysis with no reference to J-cost, φ-ladders, 8-tick periodicity, or the reality_from_one_distinction forcing chain. RS theorems (e.g., AbsoluteFloorClosure, AlexanderDuality, costAlphaLog_high_calibrated_iff) are absent from the paper's construction.","tokens_in":57002,"confidence":"high","tokens_out":154,"duration_ms":4431,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"People project human task difficulty onto AI, overestimating performance on easy tasks and underestimating it on hard ones.","keywords":["human projection","AI adoption","belief formation","task difficulty","anthropomorphism","chatbot engagement","performance evaluation"],"falsifier":"An experiment presenting AI performance data without human-comparable task labels or cues and checking whether belief updates and adoption shift from all-or-nothing to selective patterns.","tokens_in":2575,"feed_emoji":"🤖","tokens_out":628,"duration_ms":14432,"temperature":0.7,"pith_summary":"The paper establishes that humans apply their own standards of task difficulty and mistake reasonableness when assessing AI capabilities. This projection produces systematic errors in beliefs, such as over-updating after easy failures or hard successes, and treats performance as evidence of a single underlying ability. As a result, adoption decisions tend toward all-or-nothing patterns even when AI outperforms humans only selectively. A field experiment with a parenting chatbot shows that mistakes perceived as less reasonable trigger sharper drops in trust and usage. These patterns arise because people use human evaluation frameworks for AI, which misaligns expectations when AI capabilities are jagged rather than ordered like human ones.","feed_headline":"Humans project difficulty standards onto AI, skewing beliefs","feed_subtitle":"This produces overestimation on easy tasks, underestimation on hard ones, and all-or-nothing adoption decisions.","key_machinery":"Human Projection (HP), the mechanism of applying human difficulty and reasonableness standards to AI, which drives cross-task generalization and binary adoption.","core_discovery":"Human Projection (HP) is the tendency to evaluate AI using the same frameworks applied to humans, where task difficulty and the reasonableness of mistakes serve as diagnostics of overall ability. This produces overestimation on human-easy tasks, underestimation on human-hard tasks, and over-updating after easy failures and hard successes. It also induces single-index interpretations of performance, leading to all-or-nothing adoption even when superiority is task-specific, with anthropomorphic cues strengthening the effect and a field setting confirming larger trust losses from unreasonable errors.","pith_inferences":["Interfaces that present AI results without human difficulty framing could improve calibration of user expectations.","The same projection may apply to other non-human systems whose error patterns diverge from human norms.","Selective adoption could increase if users receive explicit task-by-task performance breakdowns rather than overall ability signals."],"forward_implications":["When AI performance order differs from human difficulty order, beliefs become systematically misspecified.","Removing human-like cues from AI reduces cross-task generalization and over-adoption.","Mistakes that appear unreasonable by human standards produce larger drops in trust and engagement.","Anthropomorphic design choices amplify the projection effect on adoption."],"fun_headline_variants":["Humans apply difficulty standards to AI judgments","Human Projection skews beliefs on AI task performance","Single ability index triggers all-or-nothing AI adoption","Unreasonable AI errors cause larger drops in trust","Anthropomorphic cues amplify cross-task AI generalizations"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The lab and field tasks isolate projection of human difficulty from other factors such as prior AI exposure or task-specific knowledge.","fun_headline_variants_meta":{"raw":{"variants":["Humans apply difficulty standards to AI judgments","Human Projection skews beliefs on AI task performance","Single ability index triggers all-or-nothing AI adoption","Unreasonable AI errors cause larger drops in trust","Anthropomorphic cues amplify cross-task AI generalizations"]},"model":"grok-4.3","cost_usd":0.003756,"raw_usage":{"total_tokens":1930,"prompt_tokens":640,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":37562000,"prompt_tokens_details":{"text_tokens":640,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1222,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":640,"tokens_out":68,"duration_ms":8112,"temperature":1.0,"reasoning_tokens":1222,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-24T00:30:59.219005+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment presenting AI performance data without human-comparable task labels or cues and checking whether belief updates and adoption shift from all-or-nothing to selective patterns.","supporting_citations":[],"review_version":1}