{"id":"682e0854-bf23-4e89-9fe7-7821510c31ee","arxiv_id":"2507.08230","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"From Reddit comments, the paper claims autistic people often find real human faces more uncanny than AI faces, but its participant counts, percentages, and quote attributions do not match across sections.","lead":"A study of Reddit discussions in r/autism reports that many autistic commenters feel more unease with real human faces than with AI-generated or robot faces, the opposite of the usual uncanny valley. The paper's own numbers and quotes are internally inconsistent, so the finding reads as a hypothesis for controlled testing rather than a measured result.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's headline percentages and the abstract's 'often' depend on an undefined denominator and an unstable comment-to-participant mapping; until the unit-level data are reconstructed, the central claim is unsupported.","rationale":"The reader's weakest assumption—that the analysis requires a one-to-one mapping from comments to distinct autistic participants—is precisely the load-bearing point. The paper's own text shows the unit of analysis is unstable: the same quotes are attributed to different participant IDs, the number of participants is never stated, and the comment-level theme frequencies in Section 4 do not align with the participant-level percentages in Section 3.2 and the design implications. The abstract's comparative wording ('differently', 'often') also requires a neurotypical baseline that the study did not collect, as Section 6 admits. This is not a disagreement with prior literature or a matter of theoretical taste; it is an internal consistency check that fails. A direct reconstruction of the coding dataset would settle whether the percentages are reproducible. Since the reader already recommended REJECT and my concern confirms that judgment, the verdict should remain unchanged.","tokens_in":8463,"tokens_out":3459,"duration_ms":40011,"concrete_test":"Reconstruct the unit-level dataset from the paper's coding materials: enumerate every unique comment ID, map each quote to exactly one participant ID, count unique participants and comments per theme, and recompute 75.4%, 24.6%, 14.0%, and 3.5% with the same denominator used in Section 4. If the percentages cannot be reproduced from unique participant IDs, then the prevalence claims and the abstract's 'often' are unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—75.4% atypical, 24.6% traditional, 14.0% immune, 3.5% AI-specific—and the abstract's 'often' depend on treating each retained Reddit comment as a distinct autistic participant. The paper never states the number of participants or comments used for these percentages. Section 3.1 reports 180 eligible comments, 57 substantive comments, and 42 selected for analysis, but Sections 3.2 and 4 report participant percentages with no denominator, and Section 4 also reports theme frequencies on comments (35%, 28%, 22%, 15%) that do not match the participant percentages. The unit of analysis is further destabilized by duplicate quote attribution: the same text appears as P7 and P8, P26 and P1, and P29 and P4. Because self-selected comments from r/autism are not a population sample, and because no neurotypical comparison group or controlled stimulus set is included, the comparative claims ('differently', 'inversion') and the prevalence figures cannot be derived from the data as presented. Section 6 concedes self-reported diagnosis and lack of quantitative comparison, but the abstract and design implications do not carry those caveats. This is a data-integrity issue, not merely an interpretive dispute: every one-decimal percentage in the abstract and Section 4 would be unsupported if a single commenter contributed multiple comments or if quotes were duplicated in coding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes posts and comments from Reddit's r/autism community to argue that autistic individuals may experience the uncanny valley differently, often reporting stronger discomfort with real human faces than with AI-generated ones. The author uses a hybrid thematic/content analysis of a corpus that is reduced from 180 eligible comments to 57 substantive comments to 42 selected comments, and reports participant-level percentages (75.4% atypical responses, 24.6% traditional uncanny valley effects, 14.0% complete immunity, 3.5% AI-specific discomfort) as well as comment-level theme frequencies. The paper concludes with design implications for AI avatars, robots, and virtual agents, suggesting that realism may actually reduce discomfort for many autistic users. The central qualitative observation is plausible and supported by some quoted comments, but the quantitative prevalence claims are not supported by the data as presented, owing to an undefined unit of analysis and internal inconsistencies in the reported frequencies.","tokens_in":8615,"tokens_out":5708,"duration_ms":65663,"significance":"The paper addresses a genuinely understudied question: how autistic adults describe their own experiences of the uncanny valley in naturalistic settings rather than in forced-choice laboratory tasks. The quoted comments do provide evidence that some autistic individuals report immunity to, or even inversion of, the classically described uncanny valley, and this is a valuable hypothesis-generating resource. The use of online community data is appropriate for this exploratory purpose, and the author is transparent in Section 6 about the self-reported diagnosis limitation. However, the manuscript's significance is severely undercut by the presentation of unstable comment counts as precisely quantified participant prevalence rates. If the work is reframed as a qualitative thematic study that does not claim to estimate population prevalence, it could make a modest but honest contribution; in its current form, the one-decimal percentages in the abstract and Section 4 go beyond what the data can support and risk misleading the design implications drawn from them.","major_comments":[{"comment":"The participant-level percentages (75.4%, 24.6%, 14.0%, 3.5%) are computed from an unspecified denominator and an unstable comment-to-participant mapping. Section 3.1 reports 180 eligible comments, 57 substantive comments, and 42 selected comments, but the paper never states how many distinct participants these correspond to. Worse, the same quoted text is attributed to two different participant IDs: P7 and P8 share one quote, P26 and P1 share another, and P29 and P4 share a third. If a single Reddit comment is counted as two participants, every prevalence figure in the abstract and Section 4 is invalid; if the duplicates are typographical errors, the paper must state this explicitly. Without a clear statement of the number of unique participants and a defensible comment-to-participant mapping, the abstract's 'often' and the one-decimal percentages cannot be evaluated or reproduced.","section":"§3.1 and §4"},{"comment":"The two sets of frequency numbers reported in the results are not reconciled. Section 4 states that themes appear in 35%, 28%, 22%, and 15% of comments, while Section 3.2 and Section 4.1 report participant rates of 75.4%, 24.6%, 14.0%, and 3.5%. These appear to measure different constructs (comments versus participants), and the relationship between them is never specified. The 75.4% 'atypical' participant rate cannot be derived from the quoted comments or from the four comment-level theme percentages, and the 14.0% 'complete immunity' figure has no corresponding comment-level percentage. The denominators for both sets of percentages are absent, leaving the reader unable to determine whether the claims are internally consistent. This is not a minor reporting gap; it directly undermines the quantitative claims that appear in the abstract, the introduction, and the design implications.","section":"§3.2 and §4"},{"comment":"The claim that autistic individuals experience the uncanny valley 'differently' or in an 'inverted' manner relative to neurotypical individuals is not supported by the study design, because no neurotypical comparison group and no controlled stimulus set are included. Section 6 acknowledges that the study 'does not allow for quantitative comparisons with neurotypical individuals,' yet the abstract, the introduction, and the discussion state the difference as an established finding. The within-group qualitative finding that some commenters report stronger discomfort with real faces is supported by the quotes, but the comparative framing ('differently,' 'inversion') requires a baseline that this dataset does not provide. At minimum, the paper should explicitly reframe the claims as self-reported experiences that are consistent with, but do not establish, a difference from neurotypical populations.","section":"§2, §5, and §6"},{"comment":"The 'depth and authenticity criterion' that reduces the corpus from 180 eligible comments to 57 substantive comments, and the subsequent selection of 42 comments 'deemed particularly relevant,' are not operationalized. No definitions, exclusion rules, or inter-rater reliability checks are reported, and the thematic and content coding were performed by a single author. Because the prevalence figures in Section 4 are computed from the selected 42 comments, the undocumented selection criterion directly determines the headline results. The manuscript needs a transparent audit trail, including the number of unique participants, the inclusion/exclusion decisions, and a coding table that maps each comment to its assigned theme and code.","section":"§3.1"}],"minor_comments":[{"comment":"The first sentence of Section 3.1 repeats the same phrase twice in succession: 'to explore the perceptions of AI-generated faces among individuals with autism/autistic individuals to explore the perceptions of AI-generated faces among individuals with autism/autistic individuals.'","section":"§3.1"},{"comment":"Section 4.1 identifies 'seven major themes' but then describes only five distinct themes in the following subsections; the reader cannot tell which two themes are missing or were merged.","section":"§4.1"},{"comment":"The bar chart in Figure 1 shows raw counts on a 0–16 scale, but the text reports percentages of 35%, 28%, 22%, and 15%; the relationship between the figure counts and the reported percentages is unexplained.","section":"Fig. 1 and §4"},{"comment":"The reference list contains malformed entries: [3] conflates Hunger, Müller, and Glaser/Strauss as a single reference; [10] has a truncated DOI ('https://doi.org/10.1038/s41598-') and placeholder text 'Vol. 1, No. 1, Article .'; several entries lack complete publication venue information.","section":"References"},{"comment":"The paper alternates between 'autistic individuals' and 'individuals with autism' even within the same sentence; the author should adopt one consistent usage or explain the choice.","section":"Throughout"},{"comment":"There are two typographical errors in the description of Schwind et al.: 'The used a mixed methods approach' should be 'They used a mixed methods approach,' and 'checks and jaw' should be 'cheeks and jaw.'","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The quantitative claims are the principal obstacle to publication. If the author can provide a clean de-identified dataset with unique participant IDs and a clear denominator, and can reconcile the comment-level and participant-level percentages, the paper could be revised to a defensible qualitative study. However, if the raw data cannot be reconstructed or the duplicates cannot be resolved, I would recommend rejection, because the central claims would remain unsupported. I also note the absence of a data availability statement, which is especially important for a study whose conclusions depend on a subset selection that is not fully documented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing you should know: this paper has a real contribution, but its central quantitative claims are unsupported as written. The novel angle is that it pulls naturally occurring, self-descriptive comments from r/autism about the uncanny valley and AI faces. That is genuinely missing from the lab-based literature on autism and uncanny valley, and the quotes do suggest a plausible pattern: some autistic adults report immunity to the uncanny valley, discomfort with real faces, and preference for stylized faces. That aligns with prior findings like Feng et al. and Ueyama, and it gives a testable hypothesis for future work. Credit where due: the author took a naturalistic dataset and tried to analyze it with recognized qualitative methods, and the design implications (customizable realism, stylized defaults) are reasonable heuristics, even if they outrun the evidence.\n\nThe soft spots are load-bearing, not cosmetic. The percentages in the abstract and Section 4—75.4%, 24.6%, 14.0%, 3.5%—have no stated denominator, and the paper never says how many distinct people are in the analysis. Three quotes are each attributed to two different participant IDs (P7/P8, P26/P1, P29/P4), so the unit of analysis is unstable. The theme frequencies reported in Section 4 (35%, 28%, 22%, 15%) are comment-based and do not align with the participant percentages. The screening from 180 to 57 to 42 comments uses an undefined \"depth and authenticity criterion.\" And there is no neurotypical comparison or controlled stimulus set, so any claim of \"inversion\" or of experiencing the effect \"differently\" overreaches. The limitations section honestly concedes self-reported diagnosis and lack of quantitative comparison, but the abstract and implications do not carry those caveats. None of this means the phenomenon is fabricated; the quotes are consistent with the direction. But the data as presented cannot support prevalence figures or the word \"often.\"\n\nWho is this for? HCI, social robotics, and autism technology researchers might use it as a pointer to a phenomenon worth studying, and it would make a good teaching case for content-analysis pitfalls. It is not a reliable source for design guidelines as currently written.\n\nRecommendation: it deserves peer review rather than desk rejection—the naturalistic angle is real and the field needs this conversation. But a serious referee should require a major revision: either ground every percentage in a disclosed denominator of unique participants, or drop the quantitative claims and frame the work as exploratory qualitative description. I would not cite the numbers as they stand.","headline":"Genuinely naturalistic qualitative data with a plausible hypothesis, but the headline percentages rest on an undefined denominator; worth a major revision, not acceptance as-is.","tokens_in":9310,"tokens_out":2005,"would_cite":false,"duration_ms":26000,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that autistic adults often find real human faces more uncanny than AI-generated faces, based on qualitative analysis of online autism community discussions.","keywords":["uncanny valley","autism","AI-generated faces","facial perception","neurodiversity","human-computer interaction","qualitative analysis","online community"],"falsifier":"A controlled experiment with clinically diagnosed autistic adults and matched neurotypical controls rating the same set of real and AI-generated faces would settle the claim; if autistic participants do not show a higher rate of atypical or inverted responses than controls, the reported inversion collapses. A simpler check is to count unique usernames in the original 57-comment dataset and recompute the 75.4% figure per person rather than per comment.","tokens_in":8067,"feed_emoji":"🤖","tokens_out":6616,"duration_ms":65482,"temperature":0.7,"pith_summary":"This paper sets out to learn whether autistic adults experience the uncanny valley effect in the usual way when looking at AI-generated faces. Analyzing comments from an online autism community, it finds that most commenters reported atypical responses: traditional uncanny valley discomfort was rare, some reported complete immunity, and the most common theme was feeling more unease with real human faces than with artificial ones. If that pattern holds, the usual assumption that realistic synthetic faces will unsettle autistic users is wrong for a large share of them, and designers of avatars, robots, and AI assistants should treat realism as a customizable option rather than a universal trigger.","feed_headline":"Autistic adults often find real faces more uncanny than AI faces","feed_subtitle":"For many autistic users, synthetic faces may soothe rather than unsettle.","key_machinery":"The load-bearing object is the uncanny valley hypothesis—the unease that near-human entities evoke—applied to forum comments as an interpretive frame. A hybrid thematic and content analysis sorts comments into response categories such as \"Affected When Looking at People,\" \"Affected AI-Generated,\" and \"No Effect,\" then converts comment frequencies into participant percentages. The \"inverted uncanny valley\" category carries the argument: it is the evidence that real faces, not artificial ones, are what unsettle many autistic commenters.","core_discovery":"The core discovery is an apparent inversion of the uncanny valley. In the study's online sample, 75.4% of participants showed atypical responses, 24.6% experienced the traditional effect, 14.0% reported complete immunity, and 3.5% reported discomfort specific to AI-generated content. The dominant theme, appearing in 35% of comments, was being affected when looking at people, with comments such as \"I sometimes find real people uncanny and have face recognition problems.\" The paper reads this as evidence that many autistic adults find real human faces more unsettling than synthetic ones, potentially because of differences in face processing, and concludes that AI-generated faces could reduce rather than increase anxiety for many autistic users.","pith_inferences":["If the finding generalizes, the conventional design heuristic that stylization avoids the valley should be inverted for the majority of autistic users, meaning photorealistic synthetic agents may be more accepted than cartoonish ones.","The results imply a direct experimental test: in paired comparisons, autistic adults should rate real human faces as more eerie than AI-generated faces more often than neurotypical controls do; this prediction is not tested in the paper.","The link to face-recognition difficulties suggests a mechanism worth probing: manipulating whether faces are familiar or unfamiliar, or whether identity processing is required, should modulate the inverted uncanny valley response if face recognition is the driver."],"forward_implications":["AI avatar and robot designers should expect that increasing realism will reduce discomfort for a substantial share of autistic users, not increase it.","Interfaces for autistic users should offer adjustable facial realism, from stylized to photorealistic, because 24.6% still show traditional uncanny valley sensitivity and 3.5% report AI-specific discomfort.","Social-skills training, educational avatars, and AI customer-service agents may be more comfortable than human interaction for some autistic users, which could improve engagement and reduce attrition.","A design rule to always stylize virtual characters to avoid the uncanny valley is not universal; for the majority group, realistic synthetic faces may be the preferred option."],"supporting_citations":[{"why":"Reports that children with autism do not show the uncanny valley effect, providing the direct prior evidence this study extends to adults.","marker":"[2]"},{"why":"Shows autistic adolescents disclose more to a robot than to a human interviewer, supporting the claim that artificial entities are processed differently.","marker":"[4]"},{"why":"Provides a Bayesian model of uncanny valley effects in therapeutic robots for autism, a theoretical basis for atypical responses.","marker":"[14]"},{"why":"Shows typical integration of face and body emotion cues in autism, which the discussion uses to frame altered face perception.","marker":"[1]"},{"why":"Documents persistence of the uncanny valley across embodied robots, used to motivate modality-specific responses.","marker":"[15]"},{"why":"Supplies the worked example of reflexive thematic analysis that structures the study's qualitative method.","marker":"[16]"},{"why":"The qualitative research guide underpinning the constant comparative and thematic analysis procedures.","marker":"[6]"}],"fun_headline_variants":["Autistic adults: real faces creepier than AI","Uncanny valley flips for autistic viewers","Autistic users find real faces more uncanny","Real faces spook autistic adults more than AI","AI faces may soothe autistic viewers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that each retained comment was posted by a distinct, self-identified autistic person, so that counting comments and counting people are the same operation.","fun_headline_variants_meta":{"raw":{"variants":["Autistic adults: real faces creepier than AI","Uncanny valley flips for autistic viewers","Autistic users find real faces more uncanny","Real faces spook autistic adults more than AI","AI faces may soothe autistic viewers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000489,"raw_usage":{"total_tokens":2337,"prompt_tokens":804,"completion_tokens":1533,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":420,"completion_tokens_details":{"reasoning_tokens":1464}},"tokens_in":420,"tokens_out":1533,"duration_ms":20352,"temperature":1.0,"reasoning_tokens":1464,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:25:42.790109+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment with clinically diagnosed autistic adults and matched neurotypical controls rating the same set of real and AI-generated faces would settle the claim; if autistic participants do not show a higher rate of atypical or inverted responses than controls, the reported inversion collapses. A simpler check is to count unique usernames in the original 57-comment dataset and recompute the 75.4% figure per person rather than per comment.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports that children with autism do not show the uncanny valley effect, providing the direct prior evidence this study extends to adults."},{"cited_title":"Kumazaki, Z","cited_arxiv_id":null,"evidence_quote":"Shows autistic adolescents disclose more to a robot than to a human interviewer, supporting the claim that artificial entities are processed differently."},{"cited_title":"Brewer, F","cited_arxiv_id":null,"evidence_quote":"Shows typical integration of face and body emotion cues in autism, which the discussion uses to frame altered face perception."},{"cited_title":"A., Sumioka, H., Nishio, S., Glas, D","cited_arxiv_id":null,"evidence_quote":"Documents persistence of the uncanny valley across embodied robots, used to motivate modality-specific responses."},{"cited_title":"Merriam and E.J","cited_arxiv_id":null,"evidence_quote":"The qualitative research guide underpinning the constant comparative and thematic analysis procedures."}],"review_version":1}