{"id":"42ef8352-80ee-4745-9a38-4888cd8104e1","arxiv_id":"2505.16847","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A comparative annotation study finds ChatGPT over-identifies inappropriate targeting in Reddit conversations and uncovers four new target categories beyond the standard hate speech classes.","lead":"This paper compares how human experts, crowd workers, and OpenAI's GPT-3 annotate harmful targeting language in Reddit threads from banned subreddits, and reports that ChatGPT labels far more comments as targeting than humans do. It also identifies four new target categories, such as body image and socioeconomic status, from analyzing expert disagreements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper labels the model 'ChatGPT' but reports using text-davinci-003 (GPT-3); the central over-identification finding is tied to that engine, so the claim as stated is not directly supported.","rationale":"The reader's weakest assumption concerned the reliability of AdjExpert as ground truth due to moderate expert agreement. That is a legitimate methodological concern, but the paper's Section 4.2 partially addresses it by focusing on 37 cases where experts unanimously disagree with ChatGPT, and the raw rate gap (75% vs 55%) is large enough that it likely survives some noise in the expert labels. The more load-bearing issue is the model mismatch: the central claim names 'ChatGPT' while the methods identify the model as text-davinci-003, a GPT-3 engine released before ChatGPT. This is an explicit, verifiable inconsistency inside the manuscript, not a matter of external consensus. It bears directly on the strongest claim because if the wrong model name is used, the empirical evidence supports a conclusion about a different system. The paper is still valuable and transparently reports the actual engine, so a correction or targeted re-run can resolve the issue; thus the CONDITIONAL verdict remains appropriate rather than moving to accept or reject. The missing dataset is a serious reproducibility gap but is orthogonal to the internal validity of the reported comparison.","tokens_in":13236,"tokens_out":6916,"duration_ms":55175,"concrete_test":"Check the model identifier in the annotation code or API logs to confirm the engine is text-davinci-003. Then run the exact prompts from Appendix E on ChatGPT (e.g., gpt-3.5-turbo) over the same 39 gold subthreads with temperature 0, and recompute the targeting rate and kappa vs. AdjExpert. If the over-identification pattern (75% vs 55%, kappa ~0.4) replicates, the claim can be generalized; if it does not, the paper must be restricted to text-davinci-003 or revised to explain the discrepancy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract, introduction, and Section 4.2 attribute the over-identification result to 'ChatGPT', but Section 3.5 explicitly states: 'We integrated OpenAI's GPT-3 language model into our annotation process, using the text-davinci-003 engine.' These are different systems: text-davinci-003 is an InstructGPT/GPT-3 model, while ChatGPT is a chat-tuned model (e.g., gpt-3.5-turbo). The paper's central claim as summarized by the reader, that 'ChatGPT over-identifies targeting language', therefore has an empirical referent that does not match the named model. This is not a stylistic nit: the title, abstract, and conclusion promise insights into ChatGPT's limitations, but the only LLM evaluated is text-davinci-003. If the authors intended to test ChatGPT, the prompt pipeline in Appendix E should be rerun on a ChatGPT model; if they intended to test GPT-3, the manuscript should consistently say so and avoid making claims about ChatGPT. This mismatch directly affects the scope and applicability of the central finding, even though the reported statistics may be valid for text-davinci-003.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a comparative annotation study of 'inappropriately targeting language' in Reddit subthreads from banned communities, using expert annotators, crowd annotators, and an LLM. It reports moderate expert inter-annotator agreement (comment-level Cohen's kappa 0.58), lower agreement for the LLM (0.40 against the adjudicated expert set), and a tendency of the LLM to over-identify comments as targeting (188 LLM-targeting annotations vs. 133.67 expert average). It also claims to uncover new target categories (social belief, body image, addiction, socioeconomic status) and discusses sources of expert disagreement.","tokens_in":13480,"tokens_out":3825,"duration_ms":27334,"significance":"If the central findings held for the model actually named, the study would be a useful caution against substituting LLM annotations for expert judgment in detecting targeted abuse, and it would extend prior hate-speech taxonomies. Strengths of the paper include transparent reporting of kappa scores, confusion matrices for expert pairs, and detailed annotation prompts in the appendix. The dataset, though small, is a concrete resource. However, the significance is materially weakened by the mismatch between the model used (text-davinci-003) and the model named throughout the paper (ChatGPT), and by the lack of statistical support for the new-category claims.","major_comments":[{"comment":"The paper's central claim is stated as 'ChatGPT tends to over-identify comments as targeting' (Section 4.2), but the only LLM evaluated is OpenAI's GPT-3 text-davinci-003 engine, as explicitly stated in Section 3.5 ('We integrated OpenAI's GPT-3 language model into our annotation process, using the text-davinci-003 engine'). Text-davinci-003 is not ChatGPT; the abstract, introduction, 'Main contributions,' and conclusion all attribute the results to ChatGPT. This mismatch directly affects the scope and validity of the headline finding: the reported statistics may hold for text-davinci-003, but they do not support claims about ChatGPT. The authors should either rename the model throughout to 'GPT-3 (text-davinci-003)' and adjust the title and abstract, or rerun the annotation pipeline on an actual ChatGPT model and report those results.","section":"Abstract, Section 3.5, Section 4.2"},{"comment":"The paper lists 'new targeting categories such as social belief, body image, addiction, and socioeconomic status' as a main contribution, but these categories appear only in a single sentence in Section 4.1. They are absent from the annotation schema in Section 3.2, from the prompts in Appendix E, and from Table 5's category counts. No statistics, examples, inter-annotator reliability scores, or case counts are provided for these alleged new categories. As written, this claim is unsupported. The authors need to specify how these categories were derived, how often they occur in the gold data, and whether annotators can reliably identify them.","section":"Section 4.1, Table 5"},{"comment":"The comparative evaluation treats the expert majority-vote adjudication (AdjExpert) as ground truth, but expert pairwise agreement is only moderate (comment-level kappa 0.58; Table 1), and Section 4.1 documents substantial substantive disagreements, including whether 'Sunday Gunday: Self-Defense' is targeting at all. With such a noisy reference, the raw kappa differences (0.40 for AdjExpert vs. ChatGPT, 0.58 for AdjExpert vs. AdjCrowd) and the over-identification claim in Figure 1 rest on a fragile foundation. Additionally, Table 5 reports counts (e.g., 188 vs. 133.67) without the number of comments in the gold set, so the statement that 'ChatGPT marks 75% of cases as targeting' cannot be verified. The authors should report the comment-level denominator, provide proportions with confidence intervals, and present agreement against each individual expert as well as against the majority-vote adjudication.","section":"Sections 3.4, 3.6, Table 4, Figure 1"}],"minor_comments":[{"comment":"Several crowd-annotator demographic tables contain implausible or pipeline-artifact entries: Table D.9 lists an age range of 120-130, and tables D.6-D.14 list 'CONSENT REVOKED' as a fluent language, primary language, nationality, sex, ethnicity, student status, employment status, and country of residence, with 'DATA EXPIRED' appearing as a student status and employment status. These should be cleaned or explicitly reported as missing-data placeholders.","section":"Appendix D"},{"comment":"The sentence 'These tokens, extracted from targeting comments or titles provided to ChatGPT along with their associated categories (as described in 3.5.3, offer additional context...' contains a self-reference to Section 3.5.3, but the target categories were described in Section 3.5.2. Please correct the cross-reference.","section":"Section 3.5.3"},{"comment":"Section 4.2 heavily relies on Figures 1 and 2 (confusion matrices for ChatGPT vs. AdjExpert), but these figures are not included in the manuscript text; only captions/placeholders appear. Please ensure the actual figures are embedded and that their content is consistent with the reported counts.","section":"Figures 1-4"},{"comment":"The manuscript contains several typos and inconsistencies: 'inapropriately' in the Highlights, 'CONSENTREV OKED' in Table D.8, 'Dann' in Example 4.3, and inconsistent use of 'ChatGPT' vs. 'GPT-3' in the Highlights (which correctly mentions GPT-3 text-davinci-003) and the abstract (which says ChatGPT). A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The model-identity mismatch (ChatGPT vs. text-davinci-003) is serious enough that the paper in its current form would likely draw criticism from reviewers and readers; the authors can address it either by rebranding all claims to GPT-3 or by running the study on ChatGPT. The 'new target categories' claim is currently a promissory note rather than a result. I would also recommend the editor ask for the actual confusion-matrix figures and the gold-set comment count, as these are needed for a proper evaluation of the central over-identification claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth knowing: the paper's headline finding—ChatGPT over-identifies targeting language—is actually about text-davinci-003, a GPT-3 engine. The abstract and intro say ChatGPT, but Section 3.5 names the real model. That oversells the referent. The genuinely new content is the qualitative discovery of four target categories (social belief, body image, addiction, socioeconomic status) and a careful error analysis of expert disagreements. That part is useful.\n\nWhat the paper does well: it's transparent with kappa numbers and confusion matrices, it distinguishes comment-level vs subthread-level agreement, and it gives concrete examples of why experts disagree. The comparison of expert, crowd, and LLM annotations is a reasonable setup, and the finding that the LLM over-flags (75% vs 55% targeting in gold data) is visible in the tables. The authors also acknowledge the small dataset and limited language scope in the Limitations section, which is honest.\n\nSoft spots. The ChatGPT/GPT-3 naming is the biggest one; it's not a nit because the central claim is tied to the engine. If they rerun on a chat-tuned model, fine; otherwise they should consistently say 'GPT-3 (text-davinci-003)' and stop saying ChatGPT. Second, the dataset and code aren't shipped, despite being promised as a 'benchmark data set'—that's a reproducibility gap. Third, the gold standard is only 39 subthreads, and expert agreement is moderate (comment-level kappa 0.58). That doesn't kill the comparative finding, but it does mean the ground truth is fragile. The crowd tables also look unpolished (age 120-130, 'CONSENT REVOKED' entries)—minor data-cleaning issue, but it suggests the appendix wasn't fully checked.\n\nThe new target categories come from expert interpretations of 'other', and they're plausibly defined, but the sample is small, so treat them as tentative. I'm not worried about circularity—building on their own TRAC corpus is normal, and they cite prior work.\n\nWho it's for: people working on hate speech annotation, especially comparing LLM vs human annotation. It deserves peer review because it's a legitimate empirical study with a clear method and mostly honest reporting. But the model-name issue and missing data require revision before acceptance.","headline":"A transparent but small annotation study whose central claim about ChatGPT is actually about text-davinci-003 (GPT-3); the new target categories and error analysis are the most useful parts.","tokens_in":14021,"tokens_out":2191,"would_cite":false,"duration_ms":16168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChatGPT over-identifies targeting language in Reddit discussions, labeling 75% of gold-data comments as targeting versus 55% by expert majority, while the study adds four new target categories.","keywords":["targeting language","hate speech detection","annotation","inter-annotator agreement","ChatGPT","Reddit","content moderation","target categories"],"falsifier":"Re-annotate a random sample of the 39 gold subthreads with an independent panel of, say, nine experts using the same guidelines; if the new panel's majority labels disagree with AdjExpert on more than 25% of comments, or if pairwise agreement stays near 0.58, then the measured over-identification of ChatGPT is an artifact of one particular adjudication rather than a property of the model.","tokens_in":13016,"feed_emoji":"💬","tokens_out":7663,"duration_ms":55886,"temperature":0.7,"pith_summary":"This paper tries to establish that inappropriately targeting language in online conversations is best detected by comparing three annotation routes—trained experts, a crowd, and ChatGPT—rather than trusting any single one. Using 498 Reddit conversation threads drawn from banned subreddits, the authors build an annotation framework that marks whether a comment targets someone, whether that target is inside or outside the thread, and which target category and tokens are involved. The central empirical finding is that ChatGPT over-identifies targeting: it labels 75% of the gold-data comments as targeting while expert majority vote labels only 55%, and its agreement with the experts is lower ($\\kappa = 0.40$ at comment level) than crowd-vs-expert agreement ($\\kappa = 0.58$). The authors also identify four target categories absent from standard hate speech schemas: social belief, body image, addiction, and socioeconomic status. If correct, the study implies that automated moderation tools built on current large language models will over-block neutral content unless they are calibrated against human judgment.","feed_headline":"Over-flagging: ChatGPT calls 75% of Reddit comments targeting","feed_subtitle":"An expert-crowd-ChatGPT comparison finds the AI over-labels neutral comments and reveals four new hate-target categories","key_machinery":"The load-bearing mechanism is the three-tier annotation framework: a shared guideline defines 'targeting' as language that inappropriately directs abuse at individuals or groups, distinguishes targets inside versus outside the conversation thread, and asks annotators to mark both a target category and the specific target tokens. Expert annotations are adjudicated by majority vote to form AdjExpert, crowd annotations by majority vote to form AdjCrowd, and ChatGPT receives the same task through a chain of prompts that feed its earlier outputs forward. Agreement among annotators is quantified with Cohen's $\\kappa$ at comment and subthread levels, and confusion matrices are used to separate over-identification from under-identification. The framework's role is to make the three annotation sources comparable on the same units, so that differences in labels can be attributed to the annotator type rather than to the task definition.","core_discovery":"The paper's central claim is that ChatGPT, when given the same annotation task as human annotators, cannot substitute for expert judgment in recognizing targeting language: it is more sensitive but less specific, flagging comments as targeting that experts unanimously consider neutral, such as a factual remark about airplane models, and missing cases where the target is a famous individual mentioned indirectly. The claim is supported by a comparison of three annotation sets on 39 gold subthreads adjudicated by expert majority vote: moderate expert-expert agreement ($\\kappa = 0.58$ comment-level), lower crowd agreement ($\\kappa = 0.36$), ChatGPT-vs-expert agreement of $0.40$, and a confusion-matrix pattern in which ChatGPT flags 75% of gold-data comments as targeting versus 55% for experts. The paper further claims that the standard target-category list is incomplete, and proposes social belief, body image, addiction, and socioeconomic status as categories observed in unanimous 'other' annotations.","pith_inferences":["Beyond the paper, the over-identification pattern implies that precision-oriented metrics for large language model moderation should be reported alongside sensitivity; a system that flags 75% of content will overwhelm human review queues even if it catches more true positives.","The new categories suggest a testable extension: building lexicon or few-shot classifiers for social belief, body image, addiction, and socioeconomic status and measuring whether they recover comments that current toxicity filters mark as neutral.","Because the expert gold standard itself rests on a three-person majority with only moderate pairwise agreement, a practical extension would be to collect soft labels or disagreement scores from a larger panel and treat annotation uncertainty as a first-class output.","The fact that ChatGPT interpreted titles and first comments as non-targeting hints at a positional bias; a direct follow-up could swap thread order or prompt structure to see whether over-identification shifts with position."],"forward_implications":["If ChatGPT over-identifies targeting as described, deploying current large language model annotation directly in content moderation will produce many false positives, so human review or threshold calibration remains necessary.","Because the four new categories arose from unanimous 'other' annotations, standard hate speech taxonomies miss real targeting categories; adding them should improve coverage in detection systems.","Subthread-level agreement is higher than comment-level agreement for both experts and crowd, suggesting that conversation-level aggregation is a more reliable unit for moderation decisions than isolated comments.","The gold-data subset, though small, can serve as a benchmark for comparing future automated annotation systems against human adjudication.","ChatGPT's near-zero agreement on inside-versus-outside targeting suggests that structural understanding of conversation context, not just toxicity detection, is the main bottleneck for AI annotation."],"supporting_citations":[{"why":"Supplies the Reddit subthread data set from banned subreddits that all annotators label.","marker":"[9]"},{"why":"Provides the contextual abuse dataset and thread-reconstruction procedure the sampling follows.","marker":"[8]"},{"why":"Contributes the three-level offensive/targeted/target-type schema the target categories build on.","marker":"[6]"},{"why":"Establishes the target-based analysis of hate speech whose category set the paper extends.","marker":"[5]"},{"why":"Supplies the crowd annotator pre-screening/post-screening method and span-level target annotation precedent.","marker":"[11]"},{"why":"Motivates anonymization of usernames so annotators judge content without user identity cues.","marker":"[10]"}],"fun_headline_variants":["ChatGPT over-flags Reddit comments: 75% vs experts' 55%","New hate-target categories from Reddit study: belief, body, addiction, class","AI can't replace human experts in identifying online targeting","ChatGPT mislabels neutral Reddit comments as targeting, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three experts' majority-vote adjudication (AdjExpert) is a reliable ground truth; but the experts only reach moderate pairwise agreement ($\\kappa = 0.58$ at comment level) and disagree on whole categories, so if that reference standard is unstable, the ChatGPT and crowd comparisons lose their benchmark.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT over-flags Reddit comments: 75% vs experts' 55%","New hate-target categories from Reddit study: belief, body, addiction, class","AI can't replace human experts in identifying online targeting","ChatGPT mislabels neutral Reddit comments as targeting, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1568,"prompt_tokens":892,"completion_tokens":676,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":596}},"tokens_in":508,"tokens_out":676,"duration_ms":5685,"temperature":1.0,"reasoning_tokens":596,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:53:07.347728+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of the 39 gold subthreads with an independent panel of, say, nine experts using the same guidelines; if the new panel's majority labels disagree with AdjExpert on more than 25% of comments, or if pairwise agreement stays near 0.58, then the measured over-identification of ChatGPT is an artifact of one particular adjudication rather than a property of the model.","supporting_citations":[{"cited_title":"Barbarestani, I","cited_arxiv_id":null,"evidence_quote":"Supplies the Reddit subthread data set from banned subreddits that all annotators label."},{"cited_title":"Vidgen, D","cited_arxiv_id":null,"evidence_quote":"Provides the contextual abuse dataset and thread-reconstruction procedure the sampling follows."},{"cited_title":"ElSherief, V","cited_arxiv_id":null,"evidence_quote":"Establishes the target-based analysis of hate speech whose category set the paper extends."},{"cited_title":"Barbarestani, I","cited_arxiv_id":null,"evidence_quote":"Supplies the crowd annotator pre-screening/post-screening method and span-level target annotation precedent."}],"review_version":1}