REVIEW 3 major objections 5 minor 22 references
A seven-category citation-purpose taxonomy, applied to 369 citations from NLP and computational social science papers, finds that only 11% of out-of-discipline references reflect deep engagement.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 10:09 UTC pith:PTZJFDZ3
load-bearing objection The taxonomy and annotation dataset are real contributions, but the headline engagement split is an artifact of the authors' own unvalidated mapping, so treat the main claim as conditional. the 3 major comments →
How Do We Engage with Other Disciplines? A Framework to Study Meaningful Interdisciplinary Discourse in Scholarly Publications
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that a citation-purpose taxonomy specifically designed for interdisciplinary contexts can capture the role out-of-discipline references play in shaping a paper's conceptual, methodological, and empirical claims. Applying the taxonomy to 369 agreed-upon citations from NLP+CSS publications, the authors find that deep engagement is rare: Substantiation+Basis and Basis, the two 'high engagement' purposes, together account for only 11% of citations, while Related Work and Analysis, labeled low engagement, account for 47%. They further find a statistically significant association between paper section and citation purpose, arguing that where a citation appears is predi
What carries the argument
The central object is a seven-category citation-purpose taxonomy (Substantiation+Basis, Basis, Substantiation, Use, Definition, Analysis, Related Work), developed through inductive annotation of interdisciplinary NLP papers. Each category is assigned an engagement level (High/Medium/Low) in a hand-set mapping, and section locations are also mapped to engagement levels. The taxonomy works by situating each citation in its surrounding paragraph and the paper's abstract claims, allowing the annotator to distinguish surface-level mentions from citations that ground the paper's methods or arguments.
Load-bearing premise
The load-bearing premise is the authors' hand-assigned mapping of each citation purpose and paper section to a fixed engagement level; if that mapping is wrong, the headline 11% and 47% figures are artifacts of the codebook rather than properties of the literature.
What would settle it
A concrete check: have independent experts rate the engagement of the same 369 citations without using the codebook, then compare their engagement scores to the taxonomy's labels. If expert-assessed engagement for Related Work citations is often high, or if the 11% high-engagement figure changes substantially under a different purpose-to-engagement mapping, the framework's central quantitative claim would be undermined.
If this is right
- Provides a quantitative method to assess the quality of interdisciplinary integration, moving beyond citation-count diversity metrics.
- The finding that high engagement is rare in NLP+CSS offers empirical support for critiques of shallow interdisciplinarity in this field.
- The section-purpose correlation offers a cheap proxy signal: citations in Method and Introduction sections are more likely to show deep engagement than those in Related Work sections.
- The framework can be applied to compare engagement across venues, publication years, or different disciplinary pairs.
- The low performance of automatic classifiers identifies a concrete gap for future NLP research on citation context modeling.
Where Pith is reading between the lines
- The hand-set engagement mapping makes the 11%/47% split a conclusion of the codebook; testing with alternative mappings or independent expert judgment could change the headline numbers.
- A testable extension is to apply the taxonomy to other interdisciplinary pairs (e.g., biology + computer science) to see whether low engagement rates generalize beyond NLP+CSS.
- The framework implies that institutional incentives for interdisciplinarity may be rewarding surface-level citation practices; one could test whether papers with more high-engagement citations are themselves more influential.
- The authors' finding that semantic similarity is not predictive of purpose suggests that deeper discourse parsing is needed, pointing to a research program rather than a finished measurement tool.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a seven-category citation-purpose taxonomy to study how NLP+CSS publications engage with out-of-discipline references. The authors annotate 394 in-text citations (369 with full/partial agreement), assign each purpose and section a fixed engagement level (High/Medium/Low) in Tables 3–4, and report that low-engagement purposes (Related Work + Analysis) account for 47% of citations while high-engagement purposes (Substantiation+Basis, Basis) account for 11%. They further evaluate two automatic classification approaches, finding that neither achieves sufficient performance. The paper's stated contributions are the taxonomy, the annotated dataset, the engagement analysis, and the automated classification results.
Significance. If the engagement-level construct is validated, the framework would add a qualitative dimension to bibliometric interdisciplinarity measures, addressing a recognized gap in prior citation-purpose taxonomies. The paper's strengths include a reproducible annotated dataset with a detailed codebook, explicit inter-annotator reliability reporting (α=0.6, 64% complete agreement), and an honest assessment of the limits of current LLM-based classification (including a zero F1 for the most theoretically important category, Substantiation+Basis). The main weakness is that the central engagement-level findings are derived directly from a hand-set mapping that has not been validated externally, making the empirical claims about the field's engagement depth currently unsubstantiated.
major comments (3)
- [§5.1, Tables 3–4] The headline results — low engagement (Related Work + Analysis) = 47%, high engagement = 11% — are direct arithmetic consequences of the hand-set engagement mapping in Tables 3 and 4, which assign a fixed High/Medium/Low level to each purpose and section. No external validation is provided (e.g., expert engagement ratings, sensitivity analysis, or outcome-based checks). Consequently, the claim that 'the majority of citations ... indicating only surface-level engagement' restates the authors' codebook assumptions rather than an empirical property of the NLP+CSS literature. Please either reframe these quantities as definitional to the proposed framework, or validate the mapping by comparing it to independent judgments of engagement depth (and report sensitivity to plausible alternative mappings, e.g., moving 'Related Work' or 'Analysis' to Medium).
- [§4.3 / §5.1] Section 4.3 defines engagement as composed of purpose, section, and SPECTER-based semantic relatedness, but the relatedness scores are never incorporated into the engagement levels used in Section 5.1. Figure 4 shows only that SPECTER similarity varies little across purposes; it does not test the three-component model. The claim that 'engagement cannot be reliably inferred from semantic similarity alone' is not actually demonstrated, since the three-component model is never fit or compared against a two-component model. Either integrate relatedness into the engagement computation or clearly restrict the reported analysis to the two-component (purpose + section) model.
- [§4.2, Table 1] The 'Related Work' category is defined in Table 1 as a residual catch-all ('does not fit any of the other categories'). It is the largest category (129/369 = 35%) and is labeled Low engagement in Table 4, so the 47% low-engagement finding is highly sensitive to how this residual class is interpreted. The paper reports only α=0.6 for the purpose labels; it should also report how often annotators chose 'Related Work' as a fallback and whether disagreement cases (Figure 3) concentrate in this category. Without this information, the low-engagement share may be an artifact of annotator uncertainty channeled into the residual category rather than evidence of surface-level engagement.
minor comments (5)
- [§3.2] The sentence 'We identified 26,289 of the in-text citations as out-of-discipline, for an average of out-of-discipline citations per article' is missing the average value. Also, the preceding sentence mentions '11,158 interdisciplinary' citations; clarify whether this term is synonymous with 'out-of-discipline' or a separate notion.
- [§4.2] Please define 'partial agreement' (reported as 93.91%). Without a definition, readers cannot interpret the reliability of the agreed-upon dataset.
- [Figures 4 and 5] The captions say 'Correlation between citation purpose and context relatedness' and 'Correlation between citation section and purpose,' but Figure 4 is a distribution/box plot and Figure 5 shows chi-square residuals. The captions should be reworded to describe the actual visualization.
- [Appendix A.3.7] Typo: 'he citation refers to work' should be 'The citation refers to work.'
- [Table 8] The row for 'accuracy' is misformatted ('accuracy0.327 0.327 0.327 0'); fix the spacing and column alignment.
Circularity Check
Headline engagement findings are defined by the hand-set Tables 3–4 mapping: 'high = 11%, low = 47%' restates the authors' own purpose-to-engagement assignment.
specific steps
-
self definitional
[Sec. 4.3 (Tables 3-4) and Sec. 5.1]
"Tables 3 and 4 contain the citation sections and purposes, with a qualitative estimate of the depth of engagement that they capture. ... Analysis and Related Work together, indicating low engagement, make up 47% of the analyzed citations. Conversely, purposes indicating high engagement make up only 11% of the citations."
Engagement level is defined by the authors in Table 4: Related Work and Analysis are labeled Low, while Basis and Substantiation+Basis are labeled High. The Sec. 5.1 percentages are therefore not an independent empirical finding about the literature; they are the annotated purpose distribution rescaled through this hand-set mapping. The concluding claim that 'the purpose can be predictive of the level of engagement' is tautological, since the level is assigned from the purpose. No inter-annotator reliability or external validation is reported for the engagement levels themselves (α=0.6 is for purpose labels only), and the SPECTER relatedness component is computed but never incorporated into the headline engagement counts.
full rationale
The paper builds a genuinely useful annotation dataset and tests automatic classifiers, so much of the contribution is independent and not circular. However, the central quantitative result about interdisciplinary engagement—only 11% high engagement and 47% low engagement—is not a measured property of the citations but a direct consequence of the authors' a priori Tables 3-4 mapping of citation purposes and sections to engagement levels. The counts are real annotations, but the 'High/Medium/Low' labels are definitions, not discoveries. The statement that purpose is 'predictive' of engagement reduces to the codebook, because each purpose is assigned a single engagement level in Table 4. The section-level mapping in Table 3 has the same issue: labeling Introduction 'High' and Related Work 'Low' and then reporting that most citations are in Related Work is a restatement of the mapping. The paper also relies on Leto et al. (2024), prior work by two of the same authors, for the underlying data collection, but this is a normal methodological building block rather than a circular justificatory chain; the citation-purpose taxonomy itself is adapted from external work (Abu-Jbara et al., 2013) and tested with annotation agreement. The circularity is therefore partial and localized: the framework's headline engagement split is self-definitional, while the annotated dataset, reliability measure, and classifier experiments retain independent content.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Citation purpose can be inferred from the paragraph surrounding the citation plus the paper's abstract.
- domain assumption Engagement depth is a function of citation purpose, paper section, and SPECTER relatedness, with fixed High/Medium/Low assignments.
- domain assumption Semantic Scholar field-of-study labels accurately distinguish in-discipline from out-of-discipline citations.
- domain assumption The seven taxonomy categories are exhaustive and mutually exclusive.
- domain assumption The RoBERTa track classifier trained on 3,189 abstracts identifies NLP+CSS papers sufficiently well to define the corpus.
invented entities (2)
-
Seven-category citation-purpose taxonomy (Substantiation+Basis, Basis, Substantiation, Use, Definition, Analysis, Related Work)
no independent evidence
-
Engagement level construct (High/Medium/Low) assigned to purposes and sections
no independent evidence
read the original abstract
With the rising popularity of interdisciplinary work and increasing institutional incentives in this direction, there is a growing need to understand how resulting publications incorporate ideas from multiple disciplines. Existing computational approaches, such as affiliation diversity, keywords, and citation patterns, do not account for how individual citations are used to advance the citing work. Although, in line with addressing this gap, prior studies have proposed taxonomies to classify citation purpose, these frameworks are not well-suited to interdisciplinary research and do not provide quantitative measures of citation engagement quality. To address these limitations, we propose a framework for the evaluation of citation engagement in interdisciplinary Natural Language Processing (NLP) publications. Our approach introduces a citation purpose taxonomy tailored to interdisciplinary work, supported by an annotation study. We demonstrate the utility of this framework through a thorough analysis of publications at the intersection of NLP and Computational Social Science.
Figures
Reference graph
Works this paper leans on
-
[1]
Amjad Abu-Jbara, Jefferson Ezra, and Dragomir Radev. 2013. https://aclanthology.org/N13-1067/ Purpose and polarity of citation: Towards NLP -based bibliometrics . In Proceedings of the 2013 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies , pages 596--606, Atlanta, Georgia
2013
-
[2]
Christian Baden, Christian Pipal, Martijn Schoonvelde, and Mariken A. C. G van der Velden. 2022. https://doi.org/10.1080/19312458.2021.2015574 Three Gaps in Computational Text Analysis Methods for Social Sciences : A Research Agenda . Communication Methods and Measures, 16(1):1--18
arXiv 2022
-
[3]
Lorenzo Cassi, Raphaël Champeimont, Wilfriedo Mescheba, and Élisabeth de Turckheim. 2017. https://doi.org/doi:10.1371/journal.pone.0170296 Analysing institutions interdisciplinarity by extensive use of rao-stirling diversity index . PloS
-
[4]
Arman Cohan, Waleed Ammar, Madeleine van Zuylen, and Field Cady. 2019. https://doi.org/10.18653/v1/N19-1361 Structural scaffolds for citation intent classification in scientific publications . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long a...
-
[5]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020. SPECTER: Document-level Representation Learning using Citation-informed Transformers . In ACL
2020
-
[6]
Justin Grimmer and Brandon M. Stewart. 2013. https://doi.org/10.1093/pan/mps028 Text as data: The promise and pitfalls of automatic content analysis methods for political texts . Political Analysis, 21(3):267–297
-
[7]
Suchin Gururangan, Ana Marasovi \'c , Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.acl-main.740 Don ' t stop pretraining: Adapt language models to domains and tasks . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342--8360, Online
-
[8]
David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky. 2018. https://doi.org/10.1162/tacl_a_00028 Measuring the evolution of a scientific field through citation frames . Transactions of the Association for Computational Linguistics, 6:391--406
-
[9]
Rodney Michael Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, Miles Crawford, Doug Downey, Jason Dunkelberger, Oren Etzioni, Rob Evans, Sergey Feldman, Joseph Gorney, David W. Graham, F.Q. Hu, Regan Huff, Daniel King, Sebastian Kohlmeier, Ba...
Pith/arXiv arXiv 2023
-
[10]
Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, and Kyle Lo. 2022. https://doi.org/10.18653/v1/2022.naacl-main.137 M ulti C ite: Modeling realistic citations requires moving beyond the single-sentence single-label setting . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Compu...
-
[11]
Alexandria Leto, Shamik Roy, Alexander Hoyle, Daniel Acuna, and Maria Leonor Pacheco. 2024. https://doi.org/10.18653/v1/2024.nlpcss-1.11 A first step towards measuring interdisciplinary engagement in scientific publications: A case study on NLP + CSS research . In Proceedings of the Sixth Workshop on Natural Language Processing and Computational Social Sc...
-
[12]
McCarthy and Giovanna Maria Dora Dore
Arya D. McCarthy and Giovanna Maria Dora Dore. 2023. https://doi.org/10.18653/v1/2023.acl-short.136 Theory-grounded computational text analysis . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1586--1594, Toronto, Canada
-
[13]
National Academy of Sciences , National Academy of Engineering , and Institute of Medicine . 2005. https://doi.org/10.17226/11153 Facilitating Interdisciplinary Research . The National Academies Press, Washington, DC
-
[14]
OpenAI . 2025. https://chat.openai.com/chat ChatGPT 5.2 [large language model]
2025
-
[15]
Alan L. Porter and Ismael Rafols. 2009. https://doi.org/10.1007/s11192-008-2197-2 Is science becoming more interdisciplinary? Measuring and mapping six research fields over time . Scientometrics, 81(3):719--745
-
[16]
David Pride and Petr Knoth. 2020. https://doi.org/10.1145/3383583.3398617 An authoritative approach to citation classification . In Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020, JCDL '20, page 337–340, New York, NY, USA. Association for Computing Machinery
arXiv 2020
-
[17]
Zhao Qun and Yang Menghui. 2023. https://doi.org/10.1016/j.joi.2023.101425 An efficient entropy of sum approach for measuring diversity and interdisciplinarity . Journal of Informetrics, 17(3):101425
arXiv 2023
-
[18]
Simone Teufel, Advaith Siddharthan, and Dan Tidhar. 2006. https://aclanthology.org/W06-1613/ Automatic classification of citation function . In Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, pages 103--110, Sydney, Australia
2006
-
[19]
Richard Van Noorden. 2015. https://doi.org/10.1038/525306a Interdisciplinary research by the numbers . Nature, 525:306--7
-
[20]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. https://arxiv.org/abs/19...
Pith/arXiv arXiv 2020
-
[21]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[22]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.