{"id":"842698e9-21eb-48ca-9451-19102f496dca","arxiv_id":"2506.13313","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Human readers and seven LLMs both fail to reliably distinguish real Amazon product reviews from ChatGPT-o1 generated fakes, performing near chance accuracy.","lead":"Three studies test whether people and seven large language models can tell real Amazon reviews from ChatGPT-generated fakes. Both humans and most LLMs land near chance, while relying on markedly different cues.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on 25 fake reviews from a single LLM under one hand-tuned prompt; until that generator class is shown representative of real-world LLM review fraud, 'indistinguishable' and 'review systems everywhere' outrun the evidence.","rationale":"The reader identified the same weakest assumption, and my read agrees: the external validity of the 25-review fake stimulus set is the most load-bearing point. All three studies, the headline accuracy, and the policy conclusion inherit from one generator-prompt combination. The paper's own §7.2 limits this to a single LLM and heuristic-guided prompting, which strengthens the case that the result is a proof-of-concept rather than a general law. I also considered the missing test of 50.8% against chance; that is a real reporting gap, but even a statistically significant 0.8-point effect would leave the practical 'indistinguishable' conclusion largely intact. A deeper issue is that LLM binary accuracy conflates discrimination with response bias and no AUC or confidence calibration is reported, but that would affect only the machine-side conclusion. The generation-representativeness concern is broader and affects both human and machine findings, so it is the single most load-bearing issue. Because the reader already flagged this and assigned CONDITIONAL, no verdict change is needed; the concrete test above is the check that would either retire the concern or force the authors to narrow the claim.","tokens_in":19983,"tokens_out":9455,"duration_ms":102970,"concrete_test":"Construct a matched extension set using the same Amazon pool: 25 real reviews plus 25 fakes from each of at least three other LLMs (e.g., GPT-4o, DeepSeek-R1, Llama-3-70B) under two prompt conditions—the full Table 2 heuristic prompt and a minimal 'write a product review' prompt. Run the same binary task with a fresh, adequately powered human sample (e.g., 100 participants x 30 reviews) and the same set of LLM judges, excluding each generator under test. If accuracy in any condition exceeds roughly 60% or differs from the original 50.8% by more than 5-10 percentage points, the headline claim must be restricted to the tested generator/prompt and cannot support 'review systems everywhere'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All three studies inherit their stimulus set from a single generation pipeline: ChatGPT-o1 with the hand-crafted instructions in Table 2, including explicit error injection ('Aim for an average of five mistakes for every ten reviews'), rating and helpful-vote distribution matching, sentiment calibration, and the instruction that 'a reader cannot distinguish the reviews from standard writing'. Section 7.2 concedes only a single LLM and heuristic-guided prompting were used. The central claim is conditional on these 25 fakes being representative of 'LLM-generated fake reviews' in the wild. Because the fakes were tuned to the same measured distributional features that define the real reviews (Table 3, no significant KS differences), the test evaluates a maximally matched generator rather than arbitrary LLM output. Real-world campaigns using less optimized prompts, other base models, or different product categories could produce reviews that are substantially easier (or harder) for humans and LLMs to detect. The abstract's 'review systems everywhere are now susceptible' therefore overstates a single-point result. This is a generalization threat, not an internal contradiction, and it is only partially mitigated by the acknowledgment in §7.2.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether humans and large language models can distinguish real Amazon product reviews from fake reviews generated by ChatGPT-o1 under a heuristic-guided prompting protocol. Study 1 reports that 288 human participants judged 25 real and 25 fake reviews with an overall accuracy of 50.82%, which the authors describe as chance-level. Study 2 benchmarks seven LLMs on the same 50 reviews, reporting accuracies between 35.6% and 50.0%, and claims that humans slightly outperform all tested LLMs. Study 3 correlates review features with classification difficulty and introduces the constructs of human 'scepticism bias' and LLM 'veracity bias'. The abstract concludes that humans and machines can no longer distinguish fake from real product reviews and that review systems everywhere are now susceptible to mechanised fraud.","tokens_in":20137,"tokens_out":6372,"duration_ms":64096,"significance":"If the results hold, this would be a valuable benchmark for human and LLM detection of adversarially matched product reviews, and the class-specific confusion patterns are interesting in their own right. The paper has real strengths: fake reviews were generated from explicit heuristics grounded in real-review statistics; the real and fake sets were compared on multiple distributional features; LLM judgments were repeated three times; and the data are openly available. The main contributions, however, depend on two things the current manuscript does not supply: proper inferential statistics for the chance-level and human-vs-LLM claims, and a scoping of the conclusions to the single-generator, single-prompt condition actually studied.","major_comments":[{"comment":"The headline claim that humans perform 'essentially the same as chance' at 50.82% is reported without any inferential statistic. Because accuracy is averaged over 288 participants, a one-sample test or 95% confidence interval is needed; with the reported per-participant SD of 0.08, the interval will be informative, and 'compatible with chance' is not the same as 'no better than chance' unless a formal test is reported. Please add a test that accounts for participant and review clustering (e.g., a mixed-effects logistic regression) and condition the abstract on the result.","section":"§3.4.1 / Abstract"},{"comment":"The comparison between human accuracy (50.82%) and LLM accuracy (35.6%-50.0%) is purely descriptive; no significance test or confidence interval is given, and the LLM estimates rest on only 50 reviews with three repetitions per model. The claim that humans 'slightly outperformed all tested LLMs' is therefore not established. Report a test that accounts for repeated LLM trials and participant clustering, or explicitly label the cross-system ranking as exploratory.","section":"§4.3, Table 5"},{"comment":"The abstract and title generalize beyond the evidence: the entire stimulus set uses 25 fake reviews from a single LLM (ChatGPT-o1) under one hand-crafted prompt with explicit error injection and distribution matching. Section 7.2 acknowledges the single-LLM limitation, but the manuscript still asserts as a general fact that fake product reviews are indistinguishable to humans and machines. This single-point result cannot support 'review systems everywhere are now susceptible.' The claims should either be narrowed to the specific generation pipeline or the study should be extended to multiple generators, prompts, and product categories; note also that ChatGPT-o1 itself is included among the detectors, so the measured near-chance performance may partly reflect a specific generator-detector pair.","section":"§3.2, §7.2"},{"comment":"Study 3 reports many Spearman correlations on n=50 reviews without multiple-comparison correction, so the specific heuristic conclusions (e.g., LLMs rely on word count, humans show a scepticism bias) are at risk of false positives. In addition, the KS tests in Table 3 with n=25 per group cannot 'confirm' the absence of distributional differences; the descriptive means differ substantially on word count (35.92 vs 66.44) and helpful votes (0.52 vs 1.92). Report equivalence tests or effect sizes with confidence intervals, and apply a multiple-comparison correction in Study 3.","section":"§5.2 / Table 3"}],"minor_comments":[{"comment":"The precision, recall, and F1 columns appear to be macro-averaged across the real and fake classes, but this is not stated; under the standard non-macro definition the values are inconsistent, e.g., ChatGPT-o1 with precision 0.752 and recall 0.507 would give F1 about 0.605, not 0.348. Please define the averaging scheme in the table note.","section":"Table 5"},{"comment":"The confidence scale is not defined in the measures subsection; please state the response scale (e.g., 0-100) and the exact wording used for the confidence question.","section":"§3.3.3"},{"comment":"There are typographical errors that should be corrected in revision, including 'utilsed' in Section 3, 'heuristicsh' in the Table 2 header, and 'V ADER' in Table 1.","section":"Throughout"},{"comment":"Grok-3's exclusion is justified by its constant 'real' response, but this constant-response behavior is itself a substantive finding under the paper's 'veracity bias' interpretation; consider analyzing it explicitly as a degenerate case rather than simply dropping it from Study 3.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's core result is plausible for the specific matched-generation condition, but the abstract and title overreach relative to the evidence. I would require the inferential statistics and claim-narrowing described in the major comments. I do not think additional generation conditions are strictly necessary for a revision if the claims are appropriately scoped, though they would strengthen the paper substantially. The inclusion of the generating model among the detectors should be disclosed prominently, as should the low power of the distributional matching checks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe takeaway: the core empirical result is plausible and the data are real, but the abstract overstates it. Fifty carefully selected reviews (25 real, 25 ChatGPT-o1 fakes with hand-injected errors) fool both 288 human participants and seven LLMs into near-chance performance. That is a useful benchmark, not a general law.\n\nWhat is actually new: this is the first systematic head-to-head of multiple frontier LLMs and a large human sample on the same review set, with aligned confusion matrices and a correlation-based analysis of which textual cues each group uses. The authors publish the data and code, and they do the right thing by checking that their fake reviews statistically match the real ones on length, sentiment, typos, and so on. The finding that LLMs default to \"real\" (veracity bias) and that humans are biased against polished positive reviews is a real contribution to the authenticity-judgment literature.\n\nThe soft spots are real but bounded. First, the generalization threat: 25 fakes from one model under one hand-crafted prompt—with explicit instruction to include five errors per ten reviews and to match observed distributions—are not representative of all LLM-generated reviews. Real campaigns using other models, prompts, or categories could be easier or harder to detect. The paper acknowledges this in section 7.2, but the abstract and many phrasings (\"review systems everywhere\") outrun the evidence. Second, the stats: the headline 50.8% has no confidence interval or significance test against chance at the participant level; the human-versus-LLM comparisons are not significance tested; and Study 3 runs many Spearman correlations on about 25 reviews with no correction, then excludes Grok-3 post hoc. None of this sinks the paper, but it needs to be fixed. Third, the cosine-similarity measure in Study 3 is unusual—for a review where humans split 50-50 and models split 50-50, it gives perfect alignment, which may not mean what the authors think.\n\nThis deserves peer review. The empirical structure is careful, the data are open, and the limitations are partly acknowledged. A revision that adds proper inferential tests, narrows the claims, and either defends or replaces the alignment metric would make it a solid contribution.\n\nWho is this for? Researchers in fake-review detection, consumer psychology, and AI policy. I would cite it for the benchmark, not for the strong claim.\n\nRecommendation: send to peer review. The core is sound; the decisive revisions are statistical transparency and calibrated conclusions.","headline":"A careful but narrow benchmark shows humans and LLMs both miss well-crafted fake reviews; the abstract's broad claim outruns the evidence, but the data and analysis deserve peer review.","tokens_in":20751,"tokens_out":3579,"would_cite":true,"duration_ms":36260,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"People and frontier AI models both fail to distinguish real from machine-written product reviews, three studies report.","keywords":["fake product reviews","LLM-generated text","AI content detection","consumer trust","online review authenticity","human-AI comparison","scepticism bias","veracity bias"],"falsifier":"Generate fake reviews with several different LLMs and with no injected typos, then rerun the same 50-review classification task; if human accuracy rises clearly above chance (for example, above 60%) or any LLM detector reaches high accuracy, the paper's indistinguishability claim is specific to its generation recipe rather than to LLM-generated reviews generally.","tokens_in":19733,"feed_emoji":"🛒","tokens_out":8801,"duration_ms":78961,"temperature":0.7,"pith_summary":"Three studies report that people and large language models can no longer reliably tell real product reviews from machine-written ones. In a task with 25 real and 25 ChatGPT-o1-generated reviews, 288 human judges averaged 50.8% accuracy, essentially chance, despite moderate confidence; seven frontier LLMs scored 35.6–50.0%, all at or below the human rate and most detecting only a small share of fakes. The paper explains the failures through two opposing heuristics: humans display a 'scepticism bias' that treats polished, highly positive reviews as fake, while LLMs display a 'veracity bias' that defaults to believing reviews are real and relies on shallow cues such as length. If this is right, text-based authenticity checks by readers or general-purpose AI are no longer a dependable safeguard, and online review systems need verification of purchase or authorship rather than detection.","feed_headline":"Both humans and AI fail to spot fake product reviews","feed_subtitle":"People average 50.8% accuracy—chance—and leading LLMs do no better, putting review trust at risk.","key_machinery":"The mechanism that carries the argument is the construction and comparison of the 50-review stimulus set. The authors first profile a thousand real online marketplace reviews for length, punctuation, pronoun use, past-tense verbs, idioms, mistakes, and sentiment; they convert those patterns into a prompt for ChatGPT-o1 so that the fake reviews reproduce human imperfections at a targeted rate (five mistakes per ten reviews, concentrated in mid- and low-star reviews). They then run the same 50 reviews through 288 human judges and seven LLMs under matched instructions, and use class-wise precision, recall, F1, and a cosine similarity between human and model judgment vectors to distinguish shared cues from divergent ones. Named explanatory constructs—humans' 'scepticism bias' toward too-good-to-be-true reviews and LLMs' 'veracity bias' toward accepting text as real—do the interpretive work in Study 3.","core_discovery":"At the paper's core is the finding that LLM-generated fake product reviews are indistinguishable from authentic human reviews to both people and machines. The evidence is a 50-review benchmark: 25 real reviews sampled from a public marketplace corpus and 25 fake reviews written by ChatGPT-o1 under prompts built from the corpus's observed style, including deliberate misspellings and grammar slips (about five errors per ten reviews). In the human study, accuracy was 50.82%; participants recognized 65.8% of real reviews but only 35.8% of fakes. In the LLM study, the best model matched human accuracy at 50.0% and the worst scored 35.6%, with LLMs as a group identifying only about 9% of fake reviews. The third study attributes the near-chance performance to systematic, opposed biases: human judges are skeptical of overly positive and polished reviews, while LLM judges are biased toward believing reviews are real and lean on surface richness such as length.","pith_inferences":["Editorial inference: A direct test of the generation recipe—producing fake reviews with zero injected mistakes or with several other LLMs—would show whether the near-chance result is a property of LLM text itself or partly an artifact of the authors' deliberate imperfection schedule.","Editorial inference: The paper's asymmetry (humans miss fakes, LLMs call everything real) suggests a combined human-machine pipeline would still have low recall for fakes; what is needed is calibration against known base rates rather than another detector.","Editorial inference: If the veracity bias reflects training-data priors, then detection prompts that force 'fake unless proven otherwise' or give class-balanced instructions could shift LLM behavior; the paper used a single standard prompt, so prompt sensitivity remains open.","Editorial inference: The finding implies watermarking or metadata disclosure at generation time is the more scalable intervention, because the paper's own data confirm that post-hoc text inspection by either reader or machine is at chance."],"forward_implications":["Review platforms that rely on human flagging or general LLM screening cannot expect to catch machine-written fakes; verified-purchase and provenance signals become the only dependable defense.","Consumers' scepticism bias means polished, highly positive genuine reviews will keep getting dismissed as fake, while fake negative reviews—a cheap way to damage competitors—are especially likely to be believed.","General-purpose LLMs are not a valid detector baseline for this task; any moderation pipeline using them should assume near-zero recall for fake reviews unless reweighted or fine-tuned.","In unbalanced real-world settings, where most reviews are genuine, the LLM default of 'real' will appear accurate overall while letting essentially all fakes through, making prevalence-based metrics dangerously misleading.","If indistinguishability holds across product categories, consumer research that treats online review text as ground-truth human opinion must contend with AI pollution of its data."],"supporting_citations":[{"why":"Supplies the Amazon Review 2023 dataset from which the 25 real Home & Kitchen reviews and the 1,000-review style profile were drawn.","marker":"Hou et al., 2024"},{"why":"Provides the linguistic-cue framework (verb forms, pronouns, idiosyncrasies) used to design the fake-review prompts and the feature analysis.","marker":"Banerjee et al., 2017"},{"why":"Establishes that human heuristics for recognizing AI-generated language are flawed and manipulable, the expectation the human study tests in a consumer setting.","marker":"Jakesch et al., 2023"},{"why":"Earlier benchmark for creating and detecting fake product reviews, showing humans struggle while fine-tuned detectors could work; the paper extends and contrasts with general LLMs.","marker":"Salminen et al., 2022"},{"why":"Shows people perform near chance at distinguishing GPT-4-written from human-written reviews, the direct prior result the paper pushes further by benchmarking LLM judges.","marker":"Kovács, 2024"},{"why":"Chatbot Arena leaderboard used to select the seven frontier LLM baselines evaluated in Study 2.","marker":"Chiang et al., 2024"},{"why":"Documents that watermarking fails on very short text, supporting the paper's conclusion that source-level verification, not post-hoc detection, is needed.","marker":"Dathathri et al., 2024"}],"fun_headline_variants":["Humans and AI both fail to spot fake reviews","LLM fake reviews fool people and machines alike","Chance-level detection of AI-written product reviews","Fake reviews indistinguishable to humans and LLMs","AI-generated reviews slip past both human and machine checks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result hinges on the premise that the 25 ChatGPT-o1 reviews written from these specific prompts—which deliberately add about five errors per ten reviews—stand in for the general class of LLM-generated fake reviews in the wild; other models, prompt styles, or error rates could change human and machine detection rates substantially.","fun_headline_variants_meta":{"raw":{"variants":["Humans and AI both fail to spot fake reviews","LLM fake reviews fool people and machines alike","Chance-level detection of AI-written product reviews","Fake reviews indistinguishable to humans and LLMs","AI-generated reviews slip past both human and machine checks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1276,"prompt_tokens":1013,"completion_tokens":263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":190}},"tokens_in":629,"tokens_out":263,"duration_ms":3046,"temperature":1.0,"reasoning_tokens":190,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:04:55.152539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate fake reviews with several different LLMs and with no injected typos, then rerun the same 50-review classification task; if human accuracy rises clearly above chance (for example, above 60%) or any LLM detector reaches high accuracy, the paper's indistinguishability claim is specific to its generation recipe rather than to LLM-generated reviews generally.","supporting_citations":[{"cited_title":"Y ., and Kim, J.-J","cited_arxiv_id":null,"evidence_quote":"Provides the linguistic-cue framework (verb forms, pronouns, idiosyncrasies) used to design the fake-review prompts and the feature analysis."},{"cited_title":"T., and Naaman, M","cited_arxiv_id":null,"evidence_quote":"Establishes that human heuristics for recognizing AI-generated language are flawed and manipulable, the expectation the human study tests in a consumer setting."},{"cited_title":"M., Jung, S.-g., and Jansen, B","cited_arxiv_id":null,"evidence_quote":"Earlier benchmark for creating and detecting fake product reviews, showing humans struggle while fine-tuned detectors could work; the paper extends and contrasts with general LLMs."},{"cited_title":"N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J","cited_arxiv_id":null,"evidence_quote":"Chatbot Arena leaderboard used to select the seven frontier LLM baselines evaluated in Study 2."}],"review_version":1}