{"id":"f440144b-e423-41fb-8675-e3be3cbaa4e7","arxiv_id":"2507.10559","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"NLP researchers should define cognitive terms, temper expectations, and candidly address ethical failures when speaking with the public.","lead":"Researchers in natural language processing are given a practical set of recommendations for talking to the public about language models, focused on vague terms, inflated expectations, and ethical failures. The paper is a field-specific communication guide that could make AI coverage more accurate and less hype-driven.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2 admits individual moderation is unlikely to change public perception, undercutting the central claim that the recommendations will strengthen public understanding and support.","rationale":"The reader's weakest assumption was that the three selected obstacles are the main levers shaping public understanding of NLP, with the paper's Limitations acknowledging subjective segmentation. My concern is adjacent but more specific: even granting those levers, the paper contains an internal admission in §4.2 that individual-level moderation is unlikely to change public perception. Because the recommendations are explicitly addressed to individual researchers, this admission undercuts the causal pathway for at least one of the three pillars of the central claim. The paper is honest about its scope and does not overclaim beyond a position-paper format, and the reader's CONDITIONAL verdict appropriately reflects that the effectiveness claim is untested. I do not think the verdict should change; the concern reinforces the need for either empirical testing or a softened claim, which CONDITIONAL already captures. I mark 'partial' rather than 'agree' because the reader focused on the selection of obstacles, while my concern is about an internal tension in the argument's level of analysis.","tokens_in":7871,"tokens_out":4009,"duration_ms":49487,"concrete_test":"Run a preregistered survey experiment with a representative sample of the general public. Present participants with a short researcher interview about an NLP topic in one of two conditions: (a) following the §4 recommendations (explicitly moderating expectations, citing mature applications like search and spam filtering) or (b) a control with typical hype-adjacent framing. Measure perceived credibility of NLP research, support for NLP/AI funding, and accurate understanding of limitations. If the moderation condition does not significantly increase support or understanding, the §4 recommendation fails its intended purpose, and the central claim loses one of its three pillars.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that NLP researchers adopting the recommendations will 'strengthen public understanding and encourage support for research' (Abstract). Yet §4.2 contains a direct internal admission: 'one researcher's careful moderation when speaking about their work is unlikely to change public perception of the field.' Since §1 frames the recommendations as 'especially for researchers interacting with popular media' and 'promoting their work on social media'—i.e., individual-level action—the §4 recommendation cannot deliver the claimed collective outcome without an additional coordinating mechanism or evidence of amplification. No such mechanism is specified. This is not merely a missing empirical test; it is a logical gap within the argument. The paper tries to bridge it by describing individual benefits, but those benefits are career-oriented, not the public-understanding and public-support outcomes promised in the abstract. A related weakness, acknowledged in the Limitations, is that the three selected obstacles are chosen subjectively; if the real drivers of public sentiment are economic or institutional, even the other two recommendations may miss the mark. Both issues point to the same underlying problem: the causal chain from individual researcher speech to aggregate public understanding and support is assumed, not demonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a position/guidance paper aimed at NLP researchers who engage with general audiences. It identifies three obstacles to public understanding and support of NLP research—vague cognitive terminology (§3), unreasonable expectations (§4), and ethical failures (§5)—and formulates recommendations for each, illustrated with examples from the research literature and popular press. The paper positions itself as an NLP-specific complement to existing general science-communication guidance, and it explicitly acknowledges its subjective scope in a Limitations section.","tokens_in":8065,"tokens_out":4702,"duration_ms":54238,"significance":"The paper addresses a real and timely gap: as LLM coverage in mainstream media grows, NLP researchers lack a field-specific, referenceable set of communication guidelines. Its strengths are concreteness and honesty; the examples (e.g., the r/changemyview study, hallucination terminology, and the AI-winter analogy) are current and well chosen, and the paper does not pretend to offer a comprehensive or empirically validated program. The recommendations are plausible and largely consistent with existing science-communication literature. However, the central claim that following these recommendations will 'strengthen public understanding and encourage support for research' is not demonstrated, and one recommendation section contains an explicit admission that individual-level moderation is unlikely to change public perception. If the paper is reframed as a set of reasoned heuristics rather than a mechanism with demonstrated aggregate effects, its contribution is useful; in its current form the abstract overstates the causal link.","major_comments":[{"comment":"The abstract states that the recommendations 'strengthen public understanding and encourage support for research,' but §4.2 contains the admission that 'one researcher's careful moderation when speaking about their work is unlikely to change public perception of the field.' Since §1 frames the recommendations as targeted at researchers interacting with popular media and social media, the §4 recommendation cannot deliver the promised collective outcome unless the paper specifies a coordination or amplification mechanism (e.g., professional-organization statements, journal editorial guidance, or evidence about cumulative effects of many researchers moderating consistently). None is given; the individual benefits listed are career-oriented rather than public-understanding outcomes. The paper should either add such a mechanism or revise the abstract and §6 to claim only that recommendations can improve individual researchers' communication, with aggregate effects left as an open question.","section":"§4.2 and Abstract"},{"comment":"The paper treats vague terminology, unreasonable expectations, and ethical failures as the three major obstacles to public understanding and support, but the selection is asserted rather than derived from evidence about public attitudes. The Limitations section acknowledges that 'the subjective nature of the topic space makes it difficult to segment it in a principled way,' and the paper does not engage with alternative or additional drivers such as economic insecurity, media incentives, or trust in institutions—factors that public-opinion research often links to AI attitudes. If those factors dominate, the recommendations may target the wrong levers. Because the abstract's causal claim depends on these being the relevant obstacles, the paper should either support the selection with data or reframe the contribution as an illustrative, non-exhaustive set of communication heuristics.","section":"§1 and Limitations"}],"minor_comments":[{"comment":"Several reference entries are malformed: 'Gordon V Cormack and 1 others' should use a standard author list, and the entry 'Lighthill James, Lighthill James, Sutherland Stuart, ...' contains duplicated names and should be corrected.","section":"References"},{"comment":"There is a spacing/formatting error in the sentence about 'the relevance ofunderstand, reason, and think'; it should read 'the relevance of understand, reason, and think.'","section":"§3.1"},{"comment":"The footnote explaining the paper's use of 'AI' versus 'NLP' appears well after the first use of 'AI winter' in the text; consider introducing the convention at first use in §4.1 or earlier.","section":"Footnote 1"},{"comment":"The recommendation to explain 'what should have been done differently' would benefit from at least one worked example of a researcher or institution responding constructively to an ethical failure, rather than only examples of failures.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is appropriate for a venue that publishes community-position papers, and the self-citations are used as illustrative examples rather than as the basis for the recommendations, so I do not see a citation-pattern concern. My main concern is that the abstract promises an outcome the paper itself concedes it cannot guarantee; a careful reframing of the causal claim should address this within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a clearly written position paper that translates general science-communication advice into NLP-specific recommendations. It is not an empirical study, and the abstract promises a causal outcome the paper does not demonstrate. Read it as a referenceable set of communication strategies and it is solid and useful.\n\nWhat is actually new is the field-specific material. The §3 discussion of cognitive terms — predict, read, learn, hallucinate — anchors real ambiguity debates to practical advice about defining scope. §4 draws a nice parallel between Hendler's three conditions for the 1980s AI winter and today's hype, with 'disowned' successes like search, OCR, and spam filtering used as talking points. §5 gives a concrete move most guides miss: be ready to explain what should have been done differently. These are not deep theoretical contributions, but they are well chosen and likely to help a researcher in a news interview or on social media. The related-work section properly credits general science communication sources, so the contribution is framed as an adaptation, not an invention.\n\nThe soft spots: the stress-test note lands. Section 4.2 openly says 'one researcher's careful moderation when speaking about their work is unlikely to change public perception of the field,' then the abstract claims these recommendations will 'strengthen public understanding and encourage support for research.' That is a real mismatch between the individual-level action and the collective-level claimed outcome, and the paper does not supply a coordinating mechanism or evidence of amplification. It is not a fatal flaw — the recommendations can be read as necessary rather than sufficient — but the abstract overreaches. The subjective choice of three obstacles is acknowledged in the Limitations and is defensible for a guide.\n\nThe citation pattern is fine; self-citations appear as examples of hallucination auditing and bias harm, not as evidence for the recommendations.\n\nWho this is for: NLP researchers who do public engagement and want a checklist. It will not change the field, but it can improve individual practice. It deserves serious peer review despite the abstract issue, because it is honest, coherent, and fills a small real gap.\n\nRecommendation: send it to review; ask for a revision that either softens the abstract or adds a sentence explaining how individual actions are expected to aggregate.","headline":"A useful, honest adaptation of science-communication guidance to NLP; the causal claim in the abstract is unsupported, but the practical recommendations are solid.","tokens_in":8547,"tokens_out":2540,"would_cite":false,"duration_ms":28454,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"NLP researchers can steady public support by defining terms, tempering hype, and discussing ethics openly.","keywords":["science communication","natural language processing","large language models","public understanding of AI","AI hype","research ethics","AI winters","communication recommendations"],"falsifier":"A randomized experiment in which two audiences read the same interview with an NLP researcher—one version following the paper's recommendations (defined cognitive terms, moderated expectations, frank ethics discussion) and one version using typical researcher language—and the two audiences show no measurable difference in comprehension or stated support would undercut the central claim that these communication practices strengthen public understanding and support.","tokens_in":7650,"feed_emoji":"💬","tokens_out":6962,"duration_ms":72893,"temperature":0.7,"pith_summary":"This paper argues that NLP researchers, who are suddenly in high demand as public explainers of large language models, lack field-specific guidance for talking to general audiences. It proposes a set of recommendations organized around three obstacles: vague cognitive terms like 'understanding' and 'reasoning', unreasonable expectations that invite boom-and-bust cycles like past AI winters, and ethical failures that erode public support. For each obstacle the paper pairs examples from published research and news coverage with concrete practices—define cognitive terms or replace them, cite mature applications that demonstrate long-term progress, and discuss what should have been done differently in ethically problematic work. The intended payoff is stronger public understanding of what NLP can and cannot do, and therefore more sustainable support for the field.","feed_headline":"NLP researchers: define terms, temper hype, own failures","feed_subtitle":"Field-specific advice says honest, clear communication about language models builds public understanding and support.","key_machinery":"The organizing mechanism is a three-obstacle framework that maps each threat to public understanding—vague cognitive terminology, unreasonable expectations, and ethical failures—to a corresponding communication strategy. The framework is carried by pairing concrete examples from research and news with actionable practices, and by the historical analogy to past AI winters, which supplies the reason expectations matter: hype creates short-term resources but risks a sharp collapse in support.","core_discovery":"The paper's central claim is that effective public communication by NLP researchers is a distinct skill the field should treat as part of its professional practice. Its core recommendation is that researchers should address three obstacles head-on when speaking with the public: scoping cognitive terms by explaining what 'predict', 'reason', or 'understand' mean in a given context, and preferring non-cognitive terms where possible; moderating expectations by acknowledging past AI winters and pointing to mature 'calm technologies' that have already succeeded; and candidly discussing ethical failures while explaining what should have been done differently and pointing to human-centered NLP research. The paper frames this not as a comprehensive guide but as a field-specific complement to general science-communication advice, grounded in examples from research literature and major news coverage.","pith_inferences":["Inference: Whether these three obstacles are the dominant drivers of public sentiment is untested; media incentives, economic interests, and institutional trust may matter more, which would be a natural empirical check of the paper's segmentation.","Inference: The recommendations imply concrete experiments—such as asking two audiences to read an interview with or without defined cognitive terms and comparing comprehension and support—that could turn the guidance into evidence.","Inference: If the practices work for NLP, they likely extend to other AI subfields experiencing public attention, though the paper itself restricts its claim to NLP.","Inference: The paper's own reasoning leaves open the possibility that individual researcher statements are too weak to move aggregate public opinion, which would limit the recommendations' reach."],"forward_implications":["Researchers who define cognitive terms or replace them with non-cognitive language will avoid unintended implications that 'predict', 'reason', or 'understand' carry for lay audiences.","Pointing to mature NLP applications such as search engines, spam filtering, and optical character recognition can give the public a concrete sense of progress and moderate boom-and-bust expectations.","Discussing ethical failures candidly, including what should have been done differently, can preserve public trust and signal that human-centered NLP research belongs in the community.","Adopting these practices gives individual researchers a defensible way to engage with the media that serves both their own visibility and the field's long-term credibility.","If consistently applied, these communication practices should strengthen public understanding of NLP's capabilities and limits, and thereby sustain support for research."],"supporting_citations":[{"why":"Supplies the AI-winter analysis that the expectations recommendations build on.","marker":"Hendler, 2008"},{"why":"Anchors the controversy over whether language models understand, motivating the terminology advice.","marker":"Bender and Koller, 2020"},{"why":"Documents ambiguity over cognitive terms like 'understand' in debates about large language models.","marker":"Mitchell and Krakauer, 2023"},{"why":"Provides a concrete recent ethical failure, the covert LLM chatbot on Reddit, that public discussion must reckon with.","marker":"O'Grady, 2025"},{"why":"Establishes the 52% public concern level that motivates the ethics-as-support section.","marker":"Faverio and Tyson, 2023"},{"why":"Identifies the gap in field-specific guidance by showing explainable-AI researchers struggle to connect with public audiences.","marker":"Hudson and Franklin, 2023"},{"why":"Supplies general science-communication background that the paper positions itself against.","marker":"National Academies, 2017"}],"fun_headline_variants":["Define, temper, own: NLP's public-communication fix","NLP: explain terms, curb hype, face failures","For NLP, clarity beats jargon, humility beats hype","NLP talks: scoped terms, honest limits, owned errors","Plain NLP talk: define words, calm expectations, confess"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the three chosen obstacles—vague terminology, unreasonable expectations, and ethical failures—are the main levers that shape public opinion about NLP, so that better researcher communication about them will materially increase public understanding and support.","fun_headline_variants_meta":{"raw":{"variants":["Define, temper, own: NLP's public-communication fix","NLP: explain terms, curb hype, face failures","For NLP, clarity beats jargon, humility beats hype","NLP talks: scoped terms, honest limits, owned errors","Plain NLP talk: define words, calm expectations, confess"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000268,"raw_usage":{"total_tokens":1558,"prompt_tokens":826,"completion_tokens":732,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":649}},"tokens_in":442,"tokens_out":732,"duration_ms":9775,"temperature":1.0,"reasoning_tokens":649,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:41:41.034680+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized experiment in which two audiences read the same interview with an NLP researcher—one version following the paper's recommendations (defined cognitive terms, moderated expectations, frank ethics discussion) and one version using typical researcher language—and the two audiences show no measurable difference in comprehension or stated support would undercut the central claim that these communication practices strengthen public understanding and support.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the AI-winter analysis that the expectations recommendations build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents ambiguity over cognitive terms like 'understand' in debates about large language models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a concrete recent ethical failure, the covert LLM chatbot on Reddit, that public discussion must reckon with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the 52% public concern level that motivates the ethics-as-support section."},{"cited_title":"Science Communications for Explainable Artificial Intelligence","cited_arxiv_id":"2308.16377","evidence_quote":"Identifies the gap in field-specific guidance by showing explainable-AI researchers struggle to connect with public audiences."}],"review_version":1}