REVIEW 2 major objections 4 minor 1 cited by
NLP Meets the World: Toward Improving Conversations With the Public About Natural Language Processing Research
T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read NLP researchers can steady public support by defining terms, tempering hype, and discussing ethics openly.
desk verdict A useful, honest adaptation of science-communication guidance to NLP; the causal claim in the abstract is unsupported, but the practical recommendations are solid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing mechanism is a three-obstacle framework that maps each threat to public understanding—vague cognitive terminology, unreasonable expectations, and ethical failures—to a corresponding communication strategy. The framework is carried by pairing concrete examples from research and news with actionable practices, and by the historical analogy to past AI winters, which supplies the reason expectations matter: hype creates short-term resources but risks a sharp collapse in support.
What would settle it
A randomized experiment in which two audiences read the same interview with an NLP researcher—one version following the paper's recommendations (defined cognitive terms, moderated expectations, frank ethics discussion) and one version using typical researcher language—and the two audiences show no measurable difference in comprehension or stated support would undercut the central claim that these communication practices strengthen public understanding and support.
Extended reading notes
Core claim
The paper's central claim is that effective public communication by NLP researchers is a distinct skill the field should treat as part of its professional practice. Its core recommendation is that researchers should address three obstacles head-on when speaking with the public: scoping cognitive terms by explaining what 'predict', 'reason', or 'understand' mean in a given context, and preferring non-cognitive terms where possible; moderating expectations by acknowledging past AI winters and pointing to mature 'calm technologies' that have already succeeded; and candidly discussing ethical failures while explaining what should have been done differently and pointing to human-centered NLP research. The paper frames this not as a comprehensive guide but as a field-specific complement to general science-communication advice, grounded in examples from research literature and major news coverage.
Load-bearing premise
The argument assumes that the three chosen obstacles—vague terminology, unreasonable expectations, and ethical failures—are the main levers that shape public opinion about NLP, so that better researcher communication about them will materially increase public understanding and support.
Editorial extensions
If this is right
- Researchers who define cognitive terms or replace them with non-cognitive language will avoid unintended implications that 'predict', 'reason', or 'understand' carry for lay audiences.
- Pointing to mature NLP applications such as search engines, spam filtering, and optical character recognition can give the public a concrete sense of progress and moderate boom-and-bust expectations.
- Discussing ethical failures candidly, including what should have been done differently, can preserve public trust and signal that human-centered NLP research belongs in the community.
- Adopting these practices gives individual researchers a defensible way to engage with the media that serves both their own visibility and the field's long-term credibility.
- If consistently applied, these communication practices should strengthen public understanding of NLP's capabilities and limits, and thereby sustain support for research.
Reading between the lines
- Inference: Whether these three obstacles are the dominant drivers of public sentiment is untested; media incentives, economic interests, and institutional trust may matter more, which would be a natural empirical check of the paper's segmentation.
- Inference: The recommendations imply concrete experiments—such as asking two audiences to read an interview with or without defined cognitive terms and comparing comprehension and support—that could turn the guidance into evidence.
- Inference: If the practices work for NLP, they likely extend to other AI subfields experiencing public attention, though the paper itself restricts its claim to NLP.
- Inference: The paper's own reasoning leaves open the possibility that individual researcher statements are too weak to move aggregate public opinion, which would limit the recommendations' reach.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a position/guidance paper aimed at NLP researchers who engage with general audiences. It identifies three obstacles to public understanding and support of NLP research—vague cognitive terminology (§3), unreasonable expectations (§4), and ethical failures (§5)—and formulates recommendations for each, illustrated with examples from the research literature and popular press. The paper positions itself as an NLP-specific complement to existing general science-communication guidance, and it explicitly acknowledges its subjective scope in a Limitations section.
Significance. The paper addresses a real and timely gap: as LLM coverage in mainstream media grows, NLP researchers lack a field-specific, referenceable set of communication guidelines. Its strengths are concreteness and honesty; the examples (e.g., the r/changemyview study, hallucination terminology, and the AI-winter analogy) are current and well chosen, and the paper does not pretend to offer a comprehensive or empirically validated program. The recommendations are plausible and largely consistent with existing science-communication literature. However, the central claim that following these recommendations will 'strengthen public understanding and encourage support for research' is not demonstrated, and one recommendation section contains an explicit admission that individual-level moderation is unlikely to change public perception. If the paper is reframed as a set of reasoned heuristics rather than a mechanism with demonstrated aggregate effects, its contribution is useful; in its current form the abstract overstates the causal link.
major comments (2)
- [§4.2 and Abstract] The abstract states that the recommendations 'strengthen public understanding and encourage support for research,' but §4.2 contains the admission that 'one researcher's careful moderation when speaking about their work is unlikely to change public perception of the field.' Since §1 frames the recommendations as targeted at researchers interacting with popular media and social media, the §4 recommendation cannot deliver the promised collective outcome unless the paper specifies a coordination or amplification mechanism (e.g., professional-organization statements, journal editorial guidance, or evidence about cumulative effects of many researchers moderating consistently). None is given; the individual benefits listed are career-oriented rather than public-understanding outcomes. The paper should either add such a mechanism or revise the abstract and §6 to claim only that recommendations can improve individual researchers' communication, with aggregate effects left as an open question.
- [§1 and Limitations] The paper treats vague terminology, unreasonable expectations, and ethical failures as the three major obstacles to public understanding and support, but the selection is asserted rather than derived from evidence about public attitudes. The Limitations section acknowledges that 'the subjective nature of the topic space makes it difficult to segment it in a principled way,' and the paper does not engage with alternative or additional drivers such as economic insecurity, media incentives, or trust in institutions—factors that public-opinion research often links to AI attitudes. If those factors dominate, the recommendations may target the wrong levers. Because the abstract's causal claim depends on these being the relevant obstacles, the paper should either support the selection with data or reframe the contribution as an illustrative, non-exhaustive set of communication heuristics.
minor comments (4)
- [References] Several reference entries are malformed: 'Gordon V Cormack and 1 others' should use a standard author list, and the entry 'Lighthill James, Lighthill James, Sutherland Stuart, ...' contains duplicated names and should be corrected.
- [§3.1] There is a spacing/formatting error in the sentence about 'the relevance ofunderstand, reason, and think'; it should read 'the relevance of understand, reason, and think.'
- [Footnote 1] The footnote explaining the paper's use of 'AI' versus 'NLP' appears well after the first use of 'AI winter' in the text; consider introducing the convention at first use in §4.1 or earlier.
- [§5.2] The recommendation to explain 'what should have been done differently' would benefit from at least one worked example of a researcher or institution responding constructively to an ethical failure, rather than only examples of failures.
Circularity Check
No significant circularity: the paper offers normative communication recommendations and illustrative examples, with no derivation or fitted input that reduces to its own conclusions.
full rationale
This manuscript is an opinion/guidance paper, not a derivation or empirical prediction. It identifies three obstacles to public understanding of NLP (vague terminology, unreasonable expectations, ethical failures) and offers communication recommendations. No quantity is fitted, no formal model is constructed, and no result is derived from an input in a way that is equivalent to its own conclusion. The recommendations are supported by external literature and illustrative news/research examples. Self-citations (e.g., Narayanan Venkit et al. 2024, Ghosh et al. 2024) are used as examples of hallucination-related audits and bias harms, not as load-bearing justification for the recommendations themselves. The stated limitation that the topic space is subjectively segmented, and the skeptical observation that individual moderation may not change aggregate public perception, are substantive weaknesses concerning causal efficacy and scope, but they are not circularity: the paper does not claim to have demonstrated the causal chain, and the recommendations are not defined in terms of the outcome they aim to achieve. Under the criteria of this review, no circular step can be exhibited with a specific reduction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (5)
- ad hoc to paper The primary audience-facing obstacles to public understanding of NLP are vague cognitive terminology, unreasonable expectations, and ethical failures.
- domain assumption Clearer researcher communication can materially improve public understanding and support.
- domain assumption The AI winter narrative, as told by Hendler, is a valid analogy for current LLM enthusiasm.
- domain assumption Public concern about AI, as measured by a 2023 Pew survey, is driven by identifiable ethical failures in NLP applications.
- domain assumption General science communication guidance transfers to the NLP research community.
Cite this review
Pith. "Pith review of NLP Meets the World: Toward Improving Conversations With the Public About Natural Language Processing Research." pith.science (2026). https://pith.science/paper/AN6BPYED
@misc{pith2026250710559,
author = {Pith},
title = {Pith review of: NLP Meets the World: Toward Improving Conversations With the Public About Natural Language Processing Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/AN6BPYED}},
note = {Machine review of arXiv:2507.10559}
}
read the original abstract
Recent developments in large language models (LLMs) have been accompanied by rapidly growing public interest in natural language processing (NLP). This attention is reflected by major news venues, which sometimes invite NLP researchers to share their knowledge and views with a wide audience. Recognizing the opportunities of the present, for both the research field and for individual researchers, this paper shares recommendations for communicating with a general audience about the capabilities and limitations of NLP. These recommendations cover three themes: vague terminology as an obstacle to public understanding, unreasonable expectations as obstacles to sustainable growth, and ethical failures as obstacles to continued support. Published NLP research and popular news coverage are cited to illustrate these themes with examples. The recommendations promote effective, transparent communication with the general public about NLP, in order to strengthen public understanding and encourage support for research.
Forward citations
Cited by 1 Pith paper
-
Motivations and Barriers to Communicating Software Engineering Research: Insights from Early Career Researchers
SE PhD students want to communicate for collaboration, recognition, and impact, but social anxiety, channel fragmentation, and missing guidance keep practice far behind aspiration.
Reference graph
Works this paper leans on
-
[1]
Dario Amodei. 2024. https://www.darioamodei.com/essay/machines-of-loving-grace Machines of loving grace: How AI could transform the world for the better . Online essay. https://www.darioamodei.com/essay/machines-of-loving-grace
work page 2024
-
[2]
Emily M. Bender and Alexander Koller. 2020. https://doi.org/10.18653/v1/2020.acl-main.463 Climbing towards NLU : On meaning, form, and understanding in the age of data . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185--5198, Online. Association for Computational Linguistics
-
[3]
Borhane Blili-Hamelin, Christopher Graziul, Leif Hancox-Li, Hananel Hazan, El-Mahdi El-Mhamdi, Avijit Ghosh, Katherine Heller, Jacob Metcalf, Fabricio Murai, Eryk Salvaggio, Andrew Smart, Todd Snider, Mariame Tighanimine, Talia Ringer, Margaret Mitchell, and Shiri Dori-Hacohen. 2025. https://arxiv.org/abs/2502.03689 Stop treating ` AGI ' as the north-star...
arXiv 2025
-
[4]
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language (technology) is power: A critical survey of ``bias'' in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454--5476, Online. Association for Computational Linguistics
-
[5]
Pascal Boyer. 1996. What makes anthropomorphism natural: Intuitive ontology and cultural representations. Journal of the Royal Anthropological Institute, pages 83--97
work page 1996
-
[6]
Bruce G Buchanan and Reid G Smith. 1988. Fundamentals of expert systems. Annual Review of Computer Science, 3(1):23--58
work page 1988
-
[7]
Seán Clarke, Dan Milmo, and Garry Blight. 2023. https://www.theguardian.com/technology/ng-interactive/2023/nov/01/how-ai-chatbots-like-chatgpt-or-bard-work-visual-explainer How AI chatbots like ChatGPT or Bard work – visual explainer . Accessed: 2025-06-19
work page 2023
-
[8]
Simon Coghlan, Kobi Leins, Susie Sheldrick, Marc Cheong, Piers Gooding, and Simon D'Alfonso. 2023. To chat or bot to chat: Ethical issues with using chatbots in mental health. Digital health, 9:20552076231183542
work page 2023
Show all 55 references
-
[9]
Gordon V Cormack and 1 others. 2008. Email spam filtering: A systematic review. Foundations and Trends in Information Retrieval , 1(4):335--455
2008
-
[10]
Daniel Crevier. 1993. AI: The Tumultuous History of the Search for Artificial Intelligence, 1st edition. Basic Books, New York, NY
1993
-
[11]
Alyssa D Edwards and Daniel M Shafer. 2022. When lamps have feelings: Empathy and anthropomorphism toward inanimate objects in animated films. Projections, 16(2):27--52
2022
-
[12]
Michelle Faverio and Alec Tyson. 2023. https://www.pewresearch.org/short-reads/2023/11/21/what-the-data-says-about-americans-views-of-artificial-intelligence/ What the data says about americans’ views of artificial intelligence . Pew Research Center – Short Reads
2023
-
[13]
Guillaume Fontaine, Marc-Andr \'e Maheu-Cadotte, Andr \'e ane Lavall \'e e, Tanya Mailhot, Genevi \`e ve Rouleau, Julien Bouix-Picasso, and Anne Bourbonnais. 2019. https://doi.org/10.2196/14447 Communicating science in the digital and social media ecosystem: Scoping review and...
2019 doi
-
[14]
Keith D. Foote. 2021. https://www.dataversity.net/a-brief-history-of-machine-learning/ A brief history of machine learning . Online article. https://www.dataversity.net/a-brief-history-of-machine-learning/
2021
-
[15]
Ina Fried. 2025. https://www.axios.com/2025/03/06/exclusive-russian-disinfo-floods-ai-chatbots-study-finds Exclusive: Russian disinformation floods AI chatbots, study finds . Axios. Based on a NewsGuard report
2025
-
[16]
Sourojit Ghosh, Pranav Narayanan Venkit, Sanjana Gautam, Shomir Wilson, and Aylin Caliskan. 2024. https://doi.org/10.1609/aies.v7i1.31651 Do generative AI models output harm while representing non-western cultures: Evidence from a community‑centered approach . In Proceedings o...
2024 doi
-
[17]
V.K Govindan and A.P Shivaprasad. 1990. https://doi.org/10.1016/0031-3203(90)90091-X Character recognition — a review . Pattern Recognition, 23(7):671--683
1990 doi
-
[18]
Boris D Grozdanoff. 2023. The looming shadow: Taking seriously potential existential threats brought about by artificial intelligence. Ethical Studies, pages 100--106
2023
-
[19]
Adam Hayes. 2025. https://www.investopedia.com/could-ai-be-coming-for-your-job-11749570 Is AI coming for your job? here's how to tell . Investopedia. Accessed: 2025-06-19
2025
-
[20]
Melissa Heikkilä. 2025. Margaret mitchell: artificial general intelligence is ‘just vibes and snake oil’. Financial Times. Accessed: 2025-06-19
2025
-
[21]
James Hendler. 2008. https://doi.org/10.1109/MIS.2008.20 Avoiding another AI winter . IEEE Intelligent Systems, 23(2):2--4
2008 doi
-
[22]
Mark J Hill and Simon Hengchen. 2019. Quantifying the impact of dirty ocr on historical text analysis: Eighteenth century collections online as a case study. Digital Scholarship in the Humanities, 34(4):825--843
2019
-
[23]
Rose Horowitch. 2025. https://www.theatlantic.com/economy/archive/2025/06/computer-science-bubble-ai/683242/ The computer‑science bubble is bursting . The Atlantic
2025
-
[24]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. https://doi.org/10.1145/3703155 A survey on hallucination in large language models: Principles, taxonomy, challenges, and o...
2025 doi
-
[25]
Simon Hudson and Matija Franklin. 2023. https://arxiv.org/abs/2308.16377 Science communications for explainable artificial intelligence . Preprint, arXiv:2308.16377
2023 arXiv
-
[26]
Lighthill James, Lighthill James, Sutherland Stuart, Needham Roger, and Longuet-Higgins Christopher. 1973. Artificial intelligence: A general survey. Science Research Council
1973
-
[27]
Nataliya Kosmyna, Eugene Hauptmann, Ye Tong Yuan, Jessica Situ, Xian-Hao Liao, Ashly Vivian Beresnitzky, Iris Braunstein, and Pattie Maes. 2025. https://arxiv.org/abs/2506.08872 Your brain on ChatGPT : Accumulation of cognitive debt when using an ai assistant for essay writing...
2025 arXiv
-
[28]
Kuncel and J
Nathan R. Kuncel and J. Rigdon. 2013. Communicating research findings. In Neal W. Schmitt, Scott Highhouse, and Irving B. Weiner, editors, Handbook of Psychology: Industrial and Organizational Psychology, 2nd edition, volume 12, pages 43--58. John Wiley & Sons, Inc., Hoboken, NJ
2013
-
[29]
Liddy, Woojin Paik, and Edmund S
Elizabeth D. Liddy, Woojin Paik, and Edmund S. Yu. 1994. https://doi.org/10.1145/183422.183425 Text categorization for multiple users based on semantic features from a machine-readable dictionary . ACM Trans. Inf. Syst., 12(3):278–295
1994
-
[30]
Ren \'e Marois and Jason Ivanoff. 2005. Capacity limits of information processing in the brain. Trends in cognitive sciences, 9(6):296--305
2005
-
[31]
Melanie Mitchell and David C Krakauer. 2023. The debate over understanding in AI's large language models. Proceedings of the National Academy of Sciences, 120(13):e2215907120
2023
-
[32]
Marius Mosbach, Vagrant Gautam, Tom \'a s Vergara Browne, Dietrich Klakow, and Mor Geva. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.181 From insights to actions: The impact of interpretability and analysis research on NLP . In Proceedings of the 2024 Conference on Empir...
2024 doi
-
[33]
Milton Mueller. 2024. The myth of AGI . Internet Governance Project
2024
-
[34]
Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, and Shomir Wilson. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.375 An audit on the perspectives and challenges of hallucinations in NLP . In Proceed...
2024 doi
-
[35]
National Academies of Sciences, Engineering, and Medicine . 2017. https://doi.org/10.17226/23674 Communicating science effectively: A research agenda
2017 doi
-
[36]
Claudio Novelli, Federico Casolari, Philipp Hacker, Giorgio Spedicato, and Luciano Floridi. 2024. Generative AI in EU law: Liability, privacy, intellectual property, and cybersecurity. Computer Law & Security Review, 55:106066
2024
-
[37]
Cathleen O'Grady. 2025. https://doi.org/10.1126/science.ady8074 ` U nethical' AI research on reddit under fire . Science, 388(6747):570--571. Perspective
2025 doi
-
[38]
Hirotaka Osawa, Dohjin Miyamoto, Satoshi Hase, Reina Saijo, Kentaro Fukuchi, and Yoichiro Miyake. 2022. Visions of artificial intelligence and robots in science fiction: a computational analysis. International Journal of Social Robotics, 14(10):2123--2133
2022
-
[39]
Alexandre Piquard. 2024. Chatbots are like parrots, they repeat without understanding. Le Monde (English edition). Interview with Emily M. Bender; Accessed: 2025-06-19
2024
-
[40]
Gil Press. 2025. https://www.forbes.com/sites/gilpress/2025/03/30/are-we-at-peak-ai-bubble-and-the-cusp-of-ai-moment/ Are we at peak AI bubble and the cusp of AI moment? Forbes
2025
-
[41]
Ellen Riloff and Wendy G Lehnert. 1992. Classifying texts using relevancy signatures. In AAAI, pages 329--334
1992
-
[42]
Anna Rogers and Isabelle Augenstein. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.112 What can we do to improve peer review in NLP ? In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1256--1262, Online. Association for Computational Linguistics
2020 doi
-
[43]
Reece Rogers. 2025. Ai is spreading old stereotypes to new languages and cultures. WIRED. Accessed: 2025-06-19
2025
-
[44]
Kevin Roose. 2023. How does ChatGPT really work. New York Times, 28
2023
-
[45]
A. L. Samuel. 1959. https://doi.org/10.1147/rd.33.0210 Some studies in machine learning using the game of checkers . IBM J. Res. Dev., 3(3):210–229
1959 doi
-
[46]
Candy Schwartz. 1998. Web search engines. Journal of the American Society for Information Science, 49(11):973--982
1998
-
[47]
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar. 2025. https://arxiv.org/abs/2506.06941 The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity . Preprint,...
2025 arXiv
-
[48]
Lucy Smith. 2025. Science communication for AI researchers: Introductory training. In Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI‑25), Tutorial Track. 13:00–14:00 tutorial + 14:00–15:00 drop-in session
2025
-
[49]
S.N. Srihari. 1992. https://doi.org/10.1109/5.156474 High-performance reading machines . Proceedings of the IEEE, 80(7):1120--1132
1992 doi
-
[50]
Izak Tait and Joshua Bensemann. 2024. Clipping the risks: Integrating consciousness in AGI to avoid existential crises. In International Conference on Artificial General Intelligence, pages 176--182. Springer
2024
-
[51]
Fiona J Tweedie, Sameer Singh, and David I Holmes. 1996. Neural network applications in stylometry: The federalist papers. Computers and the Humanities, 30:1--10
1996
-
[52]
Richard Waters. 2025. The diverging future of AI . Financial Times. Accessed: 2025-06-19
2025
-
[53]
Mark Weiser and John Seely Brown. 1996. Designing calm technology. PowerGrid Journal, 1(1):75--85
1996
-
[54]
Joseph Weizenbaum. 1966. Eliza—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1):36--45
1966
-
[55]
Marvin Wyrich and Stefan Wagner. 2023. https://doi.org/10.1109/ICSE-SEET58685.2023.00017 Teaching computer science students to communicate scientific findings more effectively . In Proceedings of the 45th International Conference on Software Engineering: Software Engineering E...
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.