Pith. sign in

REVIEW 4 major objections 5 minor 23 references

Emerging categories in scientific explanations

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper constructs SciExpl, an openly available human-annotated corpus of 272 scientific explanation sentences, and shows that six emergent categories merge by causal strength into a 3-class scheme with a Krippendorff alpha of 0.667.

desk verdict A small, openly available explanation sentence dataset with a plausible taxonomy, but the headline 0.667 alpha is a post-hoc fitted number, not a fair reliability estimate. read the letter →

arxiv 2505.17832 v1 pith:RDTPIWSD submitted 2025-05-23 cs.CL

classification cs.CL
keywords explanationcategorieshuman-annotateddatasetscientificexplanationsKrippendorff'salphainter-annotatoragreementexplainableAIcomputationallinguisticsbiotechnologytextmining
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to fill a gap in explainable AI: most explanation datasets are machine-generated or derived from tasks, not human explanations found in the wild. It extracts explanation sentences from open-access biotechnology and biophysics literature, and from them lets six categories emerge: causation, mechanistic causation, contrastive, correlation, functional, and pragmatic approach. After 120 crowd annotators labeled the sentences, the authors found that annotators could not reliably separate categories of similar causal strength, so they merged the six into three relations—strong, weak, and multi-path—and report a Krippendorff's alpha value of 0.667, an inter-annotator agreement measure. The openly available SciExpl corpus of 272 sentence-level explanations is the practical result, offered as a resource for studying and generating human-like explanations.

What carries the argument

The load-bearing machinery is the inductively built annotation schema. Six emergent categories are defined from data: causation, which states that one event leads to another without intermediate steps; mechanistic causation, which spells out the intervening mechanism; contrastive, which compares scenarios to explain divergent outcomes; correlation, which links variables without establishing cause; functional, which relates a trait's function to its form; and pragmatic approach, which frames explanation as choosing the convenient or effective option. The explicit-explanandum constraint (each sentence must name both what is explained and what explains it) and the later merging of six classes into three by causal strength are what make the labels usable and the 0.667 agreement attainable.

What would settle it

Run a fresh annotation study on new explanation sentences from the same biotechnology and biophysics sources, use only the 3-class scheme, and compute Krippendorff's alpha on all responses before any outlier removal or class merging; if the alpha is substantially below 0.667, the central reliability claim is not supported.

Watch

Extended reading notes

Core claim

The paper's central claim is that human explanations in scientific literature are not a jumble but fall into a small set of categories that can be derived inductively rather than imposed from theory. Applying that procedure to biotechnology and biophysics texts yields six categories—causation, mechanistic causation, contrastive, correlation, functional, pragmatic approach—each defined by the relation between the explanans and the explicit explanandum. Since annotators systematically confused categories of similar causal strength, the paper argues that the reliable signal is a coarser partition: strong relation (causation and mechanistic causation), weak relation (correlation, functional, pragmatic approach), and multi-path relation (contrastive). On this 3-class scheme, averaged inter-annotator agreement reaches a Krippendorff's alpha of 0.667, and the resulting corpus of 272 sentences, with both the 6-class and 3-class labels, is released openly.

Load-bearing premise

The central reliability claim depends on the assumption that the removal of outlier annotators and the merging of six categories into three were not chosen after inspecting the annotations; if they were, 0.667 overstates how well the scheme works on new explanations.

Editorial extensions

If this is right

  • The SciExpl corpus gives NLP researchers a set of 272 human-written scientific explanation sentences with labels at two granularities, directly usable for training explanation extraction and generation models.
  • The six-class taxonomy lets downstream work distinguish causation from mechanistic causation, a distinction that matters for explainable models that should cite mechanisms rather than mere causes.
  • The 3-class scheme's 0.667 alpha indicates that causal strength is a more reliably annotatable dimension than the finer category boundaries, guiding future annotation studies to design labels around that dimension.
  • Because every sentence was required to contain an explicit explanandum, the corpus provides self-contained explanation units that avoid cross-sentence context, making them easier to evaluate automatically.
  • The disagreement pattern itself is evidence that the boundary between causation and mechanistic causation is genuinely hard, which is useful to know for building explanation taxonomies in other domains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the 0.667 alpha is computed on sentences and labels that survived post-hoc filtering, a fresh annotation round on unseen sentences is the direct test of whether the 3-class scheme generalizes.
  • The paper leaves implicit that the 3-class labels could be used as a lightweight pretraining signal for explanation extraction, with the 6-class labels reserved for fine-grained analysis by domain experts.
  • A natural extension would be applying the same inductive annotation procedure to social-science or everyday texts, where relevance and truth enter the picture and the category distribution may shift away from causal strength.
  • One testable consequence is that contrastive explanations, grouped alone as multi-path, may be detectable automatically through cue phrases comparing alternative outcomes; a classifier using only such cues would test that hypothesis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces SciExpl, a human-annotated dataset of 272 explanation sentences extracted from biotechnology and biophysics literature, mainly from the PMC Open Access subset. The authors propose six inductively derived explanation categories (causation, mechanistic causation, contrastive, correlation, functional, pragmatic approach) and a coarser three-class scheme (strong relation, weak relation, multi-path relation). The central quantitative claim is that the three-class annotation reaches a Krippendorff's alpha of 0.667, presented as evidence of good inter-annotator agreement and dataset quality. The dataset is released through a public GitHub repository.

Significance. If the reliability claim holds, SciExpl is a useful resource for the study of human-like explanations in scientific text and for the construction of explanation-generation systems. The paper's strengths include an openly available dataset, clearly stated research questions, explicit category definitions with literature support, and a detailed annotator recruiting setup with 120 Prolific participants. However, the significance is substantially moderated by the fact that the headline agreement figure is computed after data-dependent sentence filtering and category merging on the same annotations that are used to evaluate the scheme; as reported, the alpha is best viewed as an optimistic, post-hoc estimate rather than an unbiased measure of the scheme's reliability on a fresh sample.

major comments (4)
  1. [Page 2, Classification study paragraph] The reported alpha of 0.667 is obtained only after two data-dependent decisions: excluding sentences to keep "the highest-quality 272 explanatory sentences" and merging six categories into three after observing "significant disagreement between categories of similar causal strength." Both decisions are made on the same annotation data that is then used to compute the alpha, so the figure conflates schema design with schema evaluation and is likely an upper bound on the reliability of the three-class scheme. The manuscript should report the six-class alpha, the alpha before sentence filtering, the number of sentences and annotators removed by the quality filter, and a hold-out or cross-validated estimate of the three-class alpha to support the claim that the scheme is reliable on a fresh sample.
  2. [Page 2, Classification study paragraph] The annotation description is too underspecified to support the central reliability claim. The terms "sanity checks" and "statistical outliers" are never defined, the criteria for selecting the "highest-quality" sentences are not given, the annotator instructions are not described, and the category definitions provided are literature-oriented summaries rather than an operational coding manual. Without the full annotation protocol, per-class agreement matrices, and the exact mapping rule used to merge the six classes into three, the reader cannot determine whether the filtering and merging were principled or optimized to raise agreement, and the study cannot be reproduced.
  3. [Page 2, Categories paragraph] The paper claims the categorization process was "entirely driven by the dataset, in a deductive classification originating from the text," yet each of the six categories is defined with citations to prior work (Mackie; Machamer et al.; Jacovi et al.; Mayr; Morgan & Morrison). This is not necessarily contradictory, but the manuscript should clarify how the pre-existing theoretical concepts guided the supposedly inductive derivation, and it should provide the actual definitions and examples shown to annotators. This is important because the "emerging categories" framing is part of the paper's contribution, and the current text leaves the relation between prior theory and data-driven discovery ambiguous.
  4. [Page 2, Agreement paragraph] The Krippendorff's alpha is presented as a single point estimate with no measure of uncertainty. With 272 sentences and ten evaluations per sentence, the sampling variability around 0.667 could be substantial, and the manuscript provides no confidence interval or variance estimate. Please supply a bootstrap or analytical confidence interval for the three-class alpha, and also report the six-class alpha and the alpha for the original 340 sentences before filtering, so that the reader can assess the magnitude of the improvement attributed to the merging step.
minor comments (5)
  1. [Abstract and Page 2, throughout] The name "Krippendorf" is misspelled; the correct spelling is "Krippendorff's alpha" (also in the abstract and the body text).
  2. [Page 2, Classification study paragraph] The mapping of the six categories into "strong relation," "weak relation," and "multi-path relation" is presented without explanation for the inclusion of "contrastive" under "multi-path relation." Please provide the operational rationale for this mapping, since it is a load-bearing step for the reported alpha.
  3. [Page 2, Dataset description] The paper provides no summary statistics of the dataset, such as the number of sentences per category in the six-class and three-class schemes, the distribution of sentences across biotechnology and biophysics, or the number of unique documents. Adding a small table would improve the usefulness of the resource description.
  4. [References, entry [2]] The dataset is referenced only by a GitHub URL. For archival permanence and citation tracking, the authors should provide a DOI or other persistent identifier for the dataset.
  5. [Page 2, Research questions] The phrase "in what form(s) do explanations take form" is grammatically awkward; consider rewording to "what forms do explanations take in scientific literature?"

Circularity Check

1 steps flagged · score 6.0 of 10

The reported 0.667 reliability is computed after post-hoc sentence filtering and category merging on the same annotations, making the headline agreement figure a fitted outcome rather than an independent estimate.

  1. fitted input called prediction [Methodology, paragraph beginning 'To minimize author bias...' (body text after the six category definitions)]
    "After sanity checks and the removal of statistical outliers, 10 evaluations per sentence were kept along with the highest-quality 272 explanatory sentences... significant disagreement between categories of similar causal strength was observed... After categorizing the sentences by causal strength and the number of relations, with the new categories... the average agreement between annotators improved to a value of 0.667."

    The 0.667 alpha is presented as evidence of a 'high-quality human-annotated explanation dataset.' The three-class scheme was introduced only after the authors observed 'significant disagreement' among the original six classes on the very same annotation data, and the alpha is then recomputed on those same retained sentences. The class definitions and the retained sentence set are therefore selected, at least in part, in response to the annotation outcomes. No six-class alpha, no pre-filter alpha, and no fresh/held-out annotation sample are reported, so the improvement to 0.667 cannot be distinguished from post hoc fitting. The reliability figure is thus an input-derived fit, not a prediction of how the schema would perform on new data.

full rationale

The dataset artifact itself (340 extracted explanation sentences, annotations, and the public repository) is an independent contribution, and the six category definitions are anchored in prior philosophy/NLP literature rather than derived within the paper, so the resource has content beyond the agreement number. However, the paper's central validation claim — that the 3-class notation 'achieves a 0.667 Krippendorf Alpha value' and therefore represents good annotator agreement — is undermined by the paper's own description: the 3-class mapping was chosen after observing disagreement in the 6-class annotations on the same data, and the alpha was computed after keeping only the 'highest-quality 272 explanatory sentences.' This makes the headline reliability metric a post hoc fitted summary, not an unbiased measure of schema stability. The scoring is therefore partial circularity (6), not full circularity (8-10), because the dataset construction and category content do not reduce to the alpha value. A fresh annotation study with the 3-class scheme, or reporting the 6-class and pre-filter alphas, would remove the circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

This is a corpus and annotation paper, so the ledger contains no fitted physical parameters. The key additions are the category schema and the three-class grouping, the latter chosen by hand on the basis of observed disagreement. The axioms above identify the domain assumptions on which the claim of a reliable, representative explanation dataset depends.

free parameters (1)
  • 3-class category mapping = strong: causation, mechanistic causation; weak: correlation, functional, pragmatic approach; multipath: contrastive
    The mapping was selected after observing 6-class annotator disagreements on the same sentences, and it directly determines the reported 0.667 alpha.
assumptions (4)
  • domain assumption The six proposed explanation categories are exhaustive and mutually exclusive for the corpus.
    Annotators were forced to choose among these categories for each sentence, but the paper does not test whether sentences could belong to multiple categories or none.
  • domain assumption Sentences with an explicit explanandum are representative of explanations in scientific literature.
    The paper excludes explanations that span multiple sentences or where the explained object is implicit, so the typology applies only to this subset.
  • domain assumption The filtered subset of 272 sentences and the remaining 10 annotations per sentence fairly represent the full corpus.
    The criteria for 'highest-quality' sentences and for removing outliers are not given, so the final reliability estimate depends on this unstated selection assumption.
  • domain assumption Krippendorf's alpha computed on the filtered and merged data measures the reliability of the original six-category scheme.
    The reported alpha is for the merged 3-class scheme, not the 6-class scheme, yet the dataset is presented with both classifications.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Emerging categories in scientific explanations." pith.science (2026). https://pith.science/paper/RDTPIWSD

@misc{pith2026250517832,
  author       = {Pith},
  title        = {Pith review of: Emerging categories in scientific explanations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RDTPIWSD}},
  note         = {Machine review of arXiv:2505.17832}
}
read the original abstract

Clear and effective explanations are essential for human understanding and knowledge dissemination. The scope of scientific research aiming to understand the essence of explanations has recently expanded from the social sciences to machine learning and artificial intelligence. Explanations for machine learning decisions must be impactful and human-like, and there is a lack of large-scale datasets focusing on human-like and human-generated explanations. This work aims to provide such a dataset by: extracting sentences that indicate explanations from scientific literature among various sources in the biotechnology and biophysics topic domains (e.g. PubMed's PMC Open Access subset); providing a multi-class notation derived inductively from the data; evaluating annotator consensus on the emerging categories. The sentences are organized in an openly-available dataset, with two different classifications (6-class and 3-class category annotation), and the 3-class notation achieves a 0.667 Krippendorf Alpha value.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 15 canonical work pages

  1. [1]

    www.prolific.com, last accessed 2025/04/08

    Prolific. www.prolific.com, last accessed 2025/04/08

  2. [2]

    https://github.com/gima9552/SciExplDataset

    SciExpl Dataset. https://github.com/gima9552/SciExplDataset

  3. [3]

    Information fusion99, 101805 (2023)

    Ali, S., Abuhmed, T., El-Sappagh, S., Muhammad, K., Alonso-Moral, J.M., Con- falonieri, R., Guidotti, R., Del Ser, J., Díaz-Rodríguez, N., Herrera, F.: Explainable artificial intelligence (xai): What we know and what is left to attain trustworthy artificial intelligence. Information fusion99, 101805 (2023)

  4. [4]

    The Lancet Digital Health3(11), e745–e750 (2021)

    Ghassemi, M., Oakden-Rayner, L., Beam, A.L.: The false hope of current ap- proaches to explainable artificial intelligence in health care. The Lancet Digital Health3(11), e745–e750 (2021)

  5. [5]

    part ii: Explanations

    Halpern, J.Y., Pearl, J.: Causes and explanations: A structural-model approach. part ii: Explanations. The British journal for the philosophy of science (2005)

  6. [6]

    In: Andreas, J., Narasimhan, K., Nematzadeh, A

    Hartmann, M., Sonntag, D.: A survey on improving NLP models with human explanations. In: Andreas, J., Narasimhan, K., Nematzadeh, A. (eds.) Proceed- ings of the First Workshop on Learning with Natural Language Supervision. pp. 40–47. Association for Computational Linguistics, Dublin, Ireland (May 2022). https://doi.org/10.18653/v1/2022.lnls-1.5, https://a...

  7. [7]

    Communication Methods and Measures1, 77–89 (04 2007)

    Hayes, A., krippendorff, k.: Answering the call for a standard reliability mea- sure for coding data. Communication Methods and Measures1, 77–89 (04 2007). https://doi.org/10.1080/19312450709336664

  8. [8]

    CoRRabs/2103.01378(2021), https://arxiv.org/abs/2103.01378

    Jacovi, A., Swayamdipta, S., Ravfogel, S., Elazar, Y., Choi, Y., Goldberg, Y.: Con- trastive explanations for model interpretability. CoRRabs/2103.01378(2021), https://arxiv.org/abs/2103.01378

Show all 23 references
  1. [9]

    Human Communication Research30, 411–433 (07 2004)

    krippendorff, k.: Reliability in content analysis: Some common misconceptions and recommendations. Human Communication Research30, 411–433 (07 2004). https://doi.org/10.1093/hcr/30.3.411

  2. [10]

    In: Proceedings of the 20th international conference on intelligent user interfaces

    Kulesza, T., Burnett, M., Wong, W.K., Stumpf, S.: Principles of explanatory de- bugging to personalize interactive machine learning. In: Proceedings of the 20th international conference on intelligent user interfaces. pp. 126–137 (2015)

  3. [11]

    Lewis, D.: Causal explanation (1986)

  4. [12]

    Entropy23(1), 18 (2020)

    Linardatos, P., Papastefanopoulos, V., Kotsiantis, S.: Explainable ai: A review of machine learning interpretability methods. Entropy23(1), 18 (2020)

  5. [13]

    Trends in cognitive sciences10(10), 464–470 (2006)

    Lombrozo, T.: The structure and function of explanations. Trends in cognitive sciences10(10), 464–470 (2006)

  6. [14]

    Philosophy of Science67(1), 1–25 (2000)

    Machamer, P., Darden, L., Craver, C.F.: Thinking about mechanisms. Philosophy of Science67(1), 1–25 (2000). https://doi.org/10.1086/392759

  7. [15]

    Clarendon Press, Oxford, (1974)

    Mackie, J.L.: The Cement of the Universe. Clarendon Press, Oxford, (1974)

  8. [16]

    Harvard University Press, Cambridge, MA (1988)

    Mayr, E.: Toward a New Philosophy of Biology: Observations of an Evolutionist. Harvard University Press, Cambridge, MA (1988)

  9. [17]

    In: Arguing About Science, pp

    Mill, J.S.: A system of logic. In: Arguing About Science, pp. 243–267. Routledge (2012)

  10. [18]

    Artificial intelligence267, 1–38 (2019)

    Miller, T.: Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence267, 1–38 (2019)

  11. [19]

    (eds.): Models as Mediators: Perspectives on Natural and Social Science

    Morgan, M.S., Morrison, M. (eds.): Models as Mediators: Perspectives on Natural and Social Science. Cambridge University Press (1999)

  12. [20]

    why should i trust you?

    Ribeiro, M.T., Singh, S., Guestrin, C.: " why should i trust you?" explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD interna- tional conference on knowledge discovery and data mining. pp. 1135–1144 (2016) 4 G. Magnifico, E. Barbu

  13. [21]

    Tan, C.: On the diversity and limits of human explanations. In: Carpuat, M., deMarneffe,M.C.,MezaRuiz,I.V.(eds.)Proceedingsofthe2022Conferenceofthe North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 2173–2188. Association ...

  14. [22]

    Mit Press (2012)

    Thagard, P.: The cognitive science of science: Explanation, discovery, and concep- tual change. Mit Press (2012)

  15. [23]

    Wiegreffe, S., Marasović, A.: Teach me to explain: A review of datasets for explain- able natural language processing (2021)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.