Pith. sign in

REVIEW 4 major objections 5 minor 17 references

Open problems in ageing science: A roadmap for biogerontology

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper curates 100 open problems in ageing science through community submissions and text-mining of the biomedical literature, and offers them as a roadmap for biogerontology.

desk verdict A genuinely useful community-sourced list of 100 open problems in ageing science, but the NLP-based article counts behind the 'data-informed' claims are not yet validated and should not be taken at face value. read the letter →

arxiv 2507.18602 v1 pith:E3RMO4WS submitted 2025-07-24 q-bio.OT

classification q-bio.OT
keywords ageinglongevityopenproblemsbiogerontologyresearchroadmapnaturallanguageprocessingbibliometricanalysispriorities
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to give ageing and longevity science a shared agenda by producing a curated list of one hundred open problems. The problems were gathered from 290 online and workshop submissions, reduced to 204 for analysis, and assessed with natural-language-processing tools against roughly 200,000 biomedical articles on ageing to find which questions already have an extensive literature and which are barely touched. The final list, selected by the team and published on an interactive website, is organised into eleven themes ranging from molecular mechanisms to interventions and biomarkers. The authors argue that a mix of broad long-term questions and narrow specific questions, including both well-studied and neglected topics, can steer biogerontology for the coming years.

What carries the argument

The mechanism that carries the selection process is a three-stage text-mining pipeline. Each open problem title is embedded with PubMedBERT, a language model trained on biomedical text, and matched by cosine similarity to the titles and abstracts of 200,228 articles tagged with the ageing topic; pairs scoring 0.2 or below are discarded. Surviving pairs pass through Med-CPT, a cross-encoder trained on biomedical search logs, which assigns a probability of query-article relevance, and through a natural-language-inference model that labels whether the article supports, contradicts, or is unrelated to the open problem. Only pairs with probability at least 0.8 and a supporting label are counted, and those counts feed the final manual selection, while consensus clustering of the title embeddings provides the starting grouping into eleven themes.

What would settle it

Sample a few dozen open problem-article pairs whose scores sit near the acceptance thresholds, have specialists independently label them as relevant or not, and compare those labels with the pipeline's counts; if agreement is low, or if small changes to the 0.2 and 0.8 thresholds dramatically change the ranking, the paper's prevalence claims are falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that the resulting 100-question list is a data-informed, community-grounded snapshot of what remains unknown in biogerontology, and that it can function as a roadmap for the field. Its own quantitative finding is that representation in the literature is highly uneven: the top 20 problems collectively account for 40.3% of all matched articles, with an average of 3,466 articles per problem, while the bottom 20 average only about 17 articles each. The authors read this disparity as separating entrenched questions, such as why we age and whether somatic mutations cause ageing, from emerging, often more tractable questions about biomarkers, specific therapeutic combinations, and less-studied mechanisms. Comparison with a 1977 predecessor list is used to show that the field has moved from characterising ageing towards trying to modulate it, while many older questions remain unresolved.

Load-bearing premise

The load-bearing assumption is that the automated text-matching thresholds used to count articles for each question genuinely measure how much of the ageing literature addresses that question; if the thresholds are wrong or the models misjudge relevance, the paper's well-studied versus neglected rankings lose their quantitative basis.

Editorial extensions

If this is right

  • Researchers can use the bottom-ranked problems as a catalogue of neglected, often small-scale targets that may be answerable within a few years.
  • Funding agencies and labs can align priorities around an explicit, community-generated question list rather than ad hoc choices.
  • The long-standing questions, including why we age and whether fundamental ageing processes exist, are confirmed as dominant themes that still lack consensus.
  • The comparison with the 1977 list indicates that the field's centre of gravity has shifted from describing ageing to modulating it through senolytics, partial reprogramming, and biomarker-guided trials.
  • The interactive website gives a persistent place where proposed solutions can accumulate, making the roadmap a living document.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the article-count metric could be read as a crude bibliometric map of research effort; weighting counts by year or by citation impact would test whether neglected topics are genuinely under-resourced or simply new.
  • Beyond the paper, the same pipeline could be applied to other research fields to generate comparable problem lists, though the thresholds would likely need recalibration for each field.
  • Beyond the paper, if the list becomes widely adopted it could shape what counts as a fundable or publishable ageing question, giving the subjective curation step outsized influence; publishing the rejected 104 problems with their counts would make that step auditable.
  • Beyond the paper, linking the neglected questions to grant databases or clinical trial registries would make the claimed research gaps testable rather than inferred from literature counts alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a curated list of 100 open problems in ageing and longevity science, assembled from 290 community submissions collected via a website and a three-day workshop. The 204 non-duplicate, relevant problems were then analysed with a three-stage NLP pipeline: PubMedBERT cosine similarity against PubMed articles, a Med-CPT cross-encoder relevance score, and an NLI-based support label. Article counts per problem are reported, and the top and bottom 20 problems are compared to identify well-studied versus neglected questions. The final list of 100 problems is grouped into 11 themes and published on an interactive website. The authors position the work as a successor to Strehler's 1977 list and as a data-informed roadmap for biogerontology.

Significance. If the NLP-based prevalence counts are reliable, this paper would provide a valuable, community-derived and data-informed research agenda for biogerontology, and its open code, data, and interactive website are concrete assets. The comparison with Strehler's earlier list adds historical perspective. However, the quantitative claims about which problems are well-studied or neglected are load-bearing for the 'data-informed roadmap' framing, and they rest entirely on an unvalidated NLP pipeline with arbitrary thresholds and an off-label use of NLI on interrogative titles. The paper also relies on a single-author selection step. For these reasons, the contribution is currently promising but not yet established.

major comments (4)
  1. [Methods — Literature-Driven Analysis of Open Problems Using NLP; Results — Summary of the Top and Bottom Open Problems] The article counts are produced by a three-stage filter whose thresholds (PubMedBERT cosine similarity above 0.2, Med-CPT probability at least 0.8, NLI label 'support') are presented without justification and without sensitivity analysis or manual validation. The Methods should include a gold-standard check on a random sample of problem-article pairs and a stability analysis, for example recomputing counts and top/bottom ranks at neighboring thresholds such as 0.1 and 0.3 (cosine) and 0.7 and 0.9 (Med-CPT). Without such validation, the quantitative statements in Results — 172,031 relevant pairs, the 40.3% vs 0.2% asymmetry, and the per-problem counts in Tables 1 and 2 — are not established.
  2. [Methods — Literature-Driven Analysis of Open Problems Using NLP] The NLI model is applied to open-problem titles, which are interrogative sentences, while NLI models are trained to classify entailment between declarative premise-hypothesis pairs. Assigning a 'support' label to a question is therefore off-label. This is not a minor technical issue: a specific problem such as 'How much do positive feedback loops and chain reactions contribute to ageing?' (Table 2) may be addressed in an abstract without that abstract entailing the question's declarative content, systematically deflating its count, while a broad question such as 'Why do we age?' may match many abstracts through loose semantic overlap. The authors should evaluate the NLI labels on a manually annotated sample of question-article pairs, or replace this stage with a retrieval or question-answering measure designed for interrogative queries.
  3. [Methods — Compilation of Open Problems; Methods — Grouping and Thematic Analysis of Open Problems] The Methods state that the final list of 100 open problems was selected by a single author, João Pedro de Magalhães, after seeing the NLP counts and cluster groupings. This makes the 'data-informed' claim hard to evaluate, since personal priorities and the unvalidated counts both shape the final output. The authors should report the exact selection protocol, including the number of candidate problems at each decision point and the criteria used to trade off importance, topic diversity, and article counts; ideally, a second researcher should independently select a list from the same candidates so that inter-rater agreement can be measured.
  4. [Results — Summary of the Top and Bottom Open Problems] The statement that the bottom 20 problems account for 0.2% of the dataset is based on summing problem-article pairs, but these pairs are not disjoint: a single PubMed article can match multiple open problems, especially broad ones. The aggregate percentage therefore does not represent the unique literature share of the bottom 20 problems. The authors should report the number of unique PubMed articles matched by any bottom-20 problem, and provide overlap statistics such as the Jaccard similarity between the matched article sets, before claiming that these topics occupy only 0.2% of the ageing literature.
minor comments (5)
  1. [Methods — Literature-Driven Analysis of Open Problems Using NLP] The phrase 'natural learning inference' should read 'natural language inference'.
  2. [Methods — Literature-Driven Analysis of Open Problems Using NLP] The exact PubMed query should be stated precisely, including the MeSH heading spelling and any qualifiers or mapping rules; the canonical MeSH heading is 'Aging', not 'Ageing', and the current wording is ambiguous about how the search was actually executed.
  3. [Results and Figure 1] Supplementary Tables 1 and 2 and Figure 1 are referenced but are not included in the manuscript, which prevents verification of the top/bottom lists and the theme distribution; these should be made available with the submission.
  4. [Author affiliations and main text] There are several formatting and typographical errors, including superscript affiliation markers rendered inconsistently (e.g., '¹¹', '²¹'), and the word 'purposedly' appears in the Discussion; the manuscript would benefit from a careful copyedit.
  5. [Data Availability Statement] The code repository is given as a GitHub URL, but no version, release tag, or environment specification is provided; adding a DOI or release archive and dependency list would improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a community-curated list with NLP-informed literature counts; its claims do not reduce to their inputs.

full rationale

The paper's central output is a curated list of 100 open problems assembled from community submissions and a workshop, with NLP-derived article counts used as a complementary guide. There is no derivation chain in which a result is defined in terms of another result. The NLP analysis applies external models (PubMedBert, Med-CPT, and an NLI model) to count PubMed articles matching each open problem title; these are empirical measurements, not fitted parameters renamed as predictions. The final selection is explicitly subjective: João Pedro de Magalhães selected the list 'prioritising those deemed important while ensuring a diversity of topics, using the initial groupings and article counts as a complementary guide.' The quantitative claims about top and bottom problems are outputs of the stated pipeline, and while the thresholds (cosine > 0.2, probability > 0.8, NLI support) are arbitrary and unvalidated, that is an external-validity or measurement-reliability concern, not circularity. Self-citations to Horvath and de Magalhães appear only as background references for existing debates and biomarkers, and they do not bear the weight of the list's construction or the NLP counts. No step reduces to its own input by construction, so the correct finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities. It relies on methodological choices (NLP thresholds, corpus definition) and domain assumptions about the representativeness of community input and the validity of the language models.

free parameters (4)
  • Cosine similarity threshold = 0.2
    Chosen threshold to filter unrelated open problem-article pairs before cross-encoder scoring (Methods, NLP Analysis). No justification or sensitivity analysis provided.
  • Med-CPT probability threshold = 0.8
    Chosen threshold for keeping relevant open problem-article pairs; sensitivity not explored (Methods, NLP Analysis).
  • PubMed time range = 1963-2023
    All PubMed articles under MeSH 'Ageing' in these years were included; boundary choice affects article counts (Methods, NLP Analysis).
  • MeSH term selection = Ageing
    The corpus was defined by the MeSH term 'Ageing'; other gerontology-related terms might change prevalence counts (Methods, NLP Analysis).
assumptions (4)
  • domain assumption PubMedBert embeddings and Med-CPT relevance scoring provide a valid measure of how closely an open problem is represented in the ageing literature.
    The entire prevalence analysis rests on these model outputs; no validation of the model-based counts against manual annotation is provided (Methods, NLP Analysis).
  • domain assumption The MeSH term 'Ageing' in PubMed captures the relevant ageing science literature.
    The corpus is restricted to articles indexed under this MeSH term, potentially missing relevant work in neighboring fields (Methods, NLP Analysis).
  • domain assumption Community submissions and workshop input are representative of the field's open problems.
    The initial 290 problems came from self-selected contributors and 24 workshop attendees; the paper does not assess sampling bias (Methods, Data collection of open problems).
  • domain assumption The final selection by one author (de Magalhães) reflects the most important open problems.
    Selection involved manual curation prioritizing importance and diversity; this subjectivity is disclosed but unvalidated (Methods, Grouping and Thematic Analysis of Open Problems).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Open problems in ageing science: A roadmap for biogerontology." pith.science (2026). https://pith.science/paper/E3RMO4WS

@misc{pith2026250718602,
  author       = {Pith},
  title        = {Pith review of: Open problems in ageing science: A roadmap for biogerontology},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E3RMO4WS}},
  note         = {Machine review of arXiv:2507.18602}
}
read the original abstract

The field of ageing science has gone through remarkable progress in recent decades, yet many fundamental questions remain unanswered or unexplored. Here we present a curated list of 100 open problems in ageing and longevity science. These questions were collected through community engagement and further analysed using Natural Language Processing to assess their prevalence in the literature and to identify both well-established and emerging research gaps. The final list is categorized into different topics, including molecular and cellular mechanisms of ageing, comparative biology and the use of model organisms, biomarkers, and the development of therapeutic interventions. Both long-standing questions and more recent and specific questions are featured. Our comprehensive compilation is available to the biogerontology community on our website (www.longevityknowledge.app). Overall, this work highlights current key research questions in ageing biology and offers a roadmap for fostering future progress in biogerontology.

Figures

Figures reproduced from arXiv: 2507.18602 by the authors.

Figure 1
Figure 1. presents the distribution of the 100 open problems across the 11 themes. The largest proportions were assigned to broader Ageing Mechanisms, more specific Molecular Mechanisms, and Interventions, which collectively accounted for over half of the selected problems. Themes such as Environmental and Physical Factors and Diversity in Human Ageing were less represented, which may reflect they are less explored topics in … view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

17 extracted references · 17 canonical work pages

  1. [1]

    Aging in Today’s Environment, 234 (The National Academies Press, Washington, DC, 1987)

    National Research Council (US), C.o.C.T.a.A. Aging in Today’s Environment, 234 (The National Academies Press, Washington, DC, 1987)

  2. [2]

    The new biology of ageing

    Partridge, L. The new biology of ageing. Philosophical Transactions of the Royal Society B: Biological Sciences 365, 147-154 (2010)

  3. [3]

    & Maynard, L

    McCay, C.M., Crowell, M.F. & Maynard, L. A. The Effect of Retarded Growth Upon the Length of Life Span and Upon the Ultimate Body Size: One Figure. The Journal of Nutrition 10, 63-79 (1935)

  4. [4]

    & Longo, V.D

    Fontana, L., Partridge, L. & Longo, V.D. Extending healthy life span--from yeast to humans. Science 328, 321-6 (2010)

  5. [5]

    , Plank, M

    de Magalhaes, J.P., Wuttke, D., Wood, S.H. , Plank, M. & Vora, C. Genome-environment interactions that modulate aging: powerful targets for drug discovery. Pharmacol Rev 64, 88-101 (2012)

  6. [6]

    The genetics of ageing

    Kenyon, C.J. The genetics of ageing. Nature 464, 504-12 (2010)

  7. [7]

    & Raj, K

    Horvath, S. & Raj, K. DNA methylation- based biomarkers and the epigenetic clock theory of ageing. Nat Rev Genet 19, 371-384 (2018)

  8. [8]

    Kennedy, B.K. et al. Geroscience: linking aging to chronic disease. Cell 159, 709-13 (2014)

Show all 17 references
  1. [9]

    & Kroemer, G

    Lopez-Otin, C., Blasco, M.A., Partridge, L ., Serrano, M. & Kroemer, G. Hallmarks of aging: An expanding universe. Cell 186, 243-278 (2023)

  2. [10]

    Seven knowledge gaps in modern biogerontology

    Rattan, S.I.S. Seven knowledge gaps in modern biogerontology. Biogerontology 25, 1-8 (2024)

  3. [11]

    Cohen, A.A. et al. Lack of consensus on an aging biology paradigm? A global survey reveals an agreement to disagree, and the need for an interdisciplinary framework. Mech Ageing Dev 191, 111316 (2020)

  4. [12]

    Gladyshev, V.N. et al. Disagreement on foundational principles of biological aging. PNAS Nexus 3, pgae499 (2024)

  5. [13]

    XI - Some Unexplored Av enues of Cellular Aging—Current and Future Research

    Strehler, B.L. XI - Some Unexplored Av enues of Cellular Aging—Current and Future Research. in Times, Cells, and Aging (Second Edition) (ed. Strehler, B.L.) 372-391 (Academic Press, 1977). 13 13

  6. [14]

    Gu, Y. et al. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans. Comput. Healthcare 3, Article 2 (2021)

  7. [15]

    Jin, Q. et al. MedCPT: Contrastive Pre-trained Transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval. Bioinformatics 39(2023)

  8. [16]

    Distinguishing between driver and passenger mechanisms of aging

    de Magalhães, J.P. Distinguishing between driver and passenger mechanisms of aging. Nat Genet 56, 204-211 (2024)

  9. [17]

    Lu, A.T. et al. Universal DNA methylation age across mammalian tissues. Nat Aging 3, 1144-1166 (2023)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.