Pith. sign in

REVIEW 5 major objections 4 minor 32 references

This paper aims to make LLM-generated biological hypotheses auditable by binding them to stable evidence IDs and two quantitative tests, showing in a 9-week Cell Painting study of low-dose radiation that every citation was valid while morph

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:15 UTC pith:N4WVI2G4

load-bearing objection A genuinely useful auditing framework for LLM-generated morphology hypotheses, but V2 needs a null baseline and direction-aware scoring before the headline compatibility claim holds. the 5 major comments →

arxiv 2607.19415 v1 pith:N4WVI2G4 submitted 2026-07-17 q-bio.QM cs.AIcs.CL

Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology

classification q-bio.QM cs.AIcs.CL
keywords Cell Paintingmorphological profilingretrieval-augmented generationLLM auditinghypothesis generationlow-dose radiationcitation validityproxy-based evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that an evaluation-first, retrieval-augmented pipeline can turn longitudinal Cell Painting morphology data into biological hypotheses that are quantitatively auditable, even when ground-truth mechanisms are unknown. Applied to a 9-week RPE-1 time course across five low-dose-rate radiation conditions, the framework produced 136 week-by-dose hypotheses, all of which passed a citation-validity check (V1) that verifies every cited evidence ID exists in the prompt. A second test (V2) uses a hand-specified proxy table to ask whether each predicted biological process is compatible with the most changed morphology features; scores rose with dose rate and correlated positively with an independent, morphology-only PCA drift summary. The authors present their proxy-based evaluation as a transparent diagnostic rather than a mechanistic benchmark, explicitly noting the absence of ground-truth mechanism labels as a limitation.

Core claim

The central claim is that an LLM interpretation of weak, chronic perturbation effects can be made scientifically usable if every hypothesis is a structured artifact carrying stable evidence identifiers, and if two quantitative audits check grounding and biological consistency. V1, a set-inclusion test, found zero invalid evidence references across all 136 hypotheses, meaning no cited observation, neighbor, or literature snippet fell outside the prompt payload. V2, which maps 11 controlled process labels to expected morphology feature shifts via an explicit proxy table, produced global mean Hit@10 of 0.4335, with best-per-dose Hit@10 increasing from 0.236 at the lowest dose rate to 0.889 at t

What carries the argument

The load-bearing object is the prompt payload built from week-matched treated–control morphology deltas bound to retrieved evidence through stable evidence identifiers: observation IDs, retrieved-neighbor IDs, literature-snippet IDs, and pathway IDs. These identifiers form the closed reference set that V1 checks for citation validity. The second mechanism is the V2 proxy-morphology compatibility score, which uses a hand-authored table associating each of the 11 controlled process labels with expected shifts in a reduced set of morphology features, then computes Hit@10 and a magnitude-weighted proxy score over the top changed features. Hierarchical integration preserves these identifiers acro

Load-bearing premise

The claim that V2 measures biological compatibility rests on a hand-written table that maps each process label to a list of morphology features and assumes that any of those features appearing among the top changed features counts as support, regardless of whether the feature moved in the direction the table specifies.

What would settle it

Recompute V2 with a signed version of Hit@10 that only counts proxy features whose observed delta direction matches the direction code in the proxy table, then re-examine the dose-response trend; if the trend flattens or reverses, the reported compatibility is an artifact of directionless set membership. Alternatively, run the framework on a spike-in dataset with known mechanism labels (e.g., a senescence inducer and an apoptosis inducer) and check whether the highest-scoring V2 labels match the known ground truth.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any future hypothesis that cites an out-of-payload evidence ID would be caught automatically by V1, making reference hallucination a detectable rather than silent failure mode.
  • V2 can flag biologically contradictory labels (for example, predicting apoptosis when the morphology signature shows massive cell enlargement) without needing external mechanism labels.
  • The positive dose-response trend in V2 implies that weak, chronic perturbations produce genuinely harder interpretation problems, so lower confidence is warranted when delta magnitudes are small.
  • The low-dose adaptive phenotype (metabolic reprogramming and proteostatic stress at 0.003–0.3 mGy/hr) is a concrete, testable prediction for follow-up molecular experiments.
  • Because evidence identifiers propagate through dose-time and global summaries, the audit applies not only to atomic hypotheses but also to higher-level integrative claims.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The direction codes in the proxy table are not used in Hit@10 or the weighted proxy score, so two features that move in opposite directions can both count as hits; a signed variant of V2 that requires direction matching would be a sharper test of biological compatibility.
  • The perfect V1 result is partly by construction, since the generation pipeline rejects and retries outputs with invalid references; removing the retry loop or enlarging the evidence pool would reveal whether grounding still holds under less constrained decoding.
  • The week-by-week V2–drift correlation is modest and its bootstrap confidence interval includes values near zero; a spike-in experiment with known mechanisms (for example, a senescence inducer versus an apoptosis inducer) would separate genuine proxy alignment from label-matching artifacts.
  • The same stable-ID audit structure should transfer to other longitudinal high-content assays and to orthogonal modalities like transcriptomics or proteomics, but the paper does not demonstrate that transfer; that would be a natural next step rather than an established result.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper proposes an evaluation-first, retrieval-augmented framework for interpreting longitudinal Cell Painting morphology data, applied to a 9-week RPE-1 time course across five low-dose-rate radiation conditions. The pipeline computes week-matched treated–control morphology deltas, retrieves perturbation neighbors (JUMP), pathway context (Reactome), and literature snippets (Europe PMC), and uses an LLM to generate structured hypotheses at week×dose resolution, which are then hierarchically summarized with stable evidence IDs. Two auditing metrics are introduced: V1 checks that cited evidence IDs exist in the prompt payload, and V2 measures overlap between the predicted process label and a hand-authored proxy set of morphological features (Hit@10 and weighted WP). The authors report V1 with no invalid references and V2 scores that increase with dose rate and correlate positively with a morphology drift summary. They explicitly acknowledge the proxy-based nature of V2 and the absence of ground-truth labels, and they release code and data.

Significance. If the framework works as claimed, it is a valuable contribution: it addresses a real need for auditable LLM interpretation in high-content imaging, offers a concrete provenance-preserving prompt schema, and proposes a quantitative compatibility check in a setting with no ground-truth mechanism labels. The hierarchical reasoning design and stable evidence IDs are well-motivated and useful beyond this specific dataset. The authors are unusually transparent about limitations, which strengthens the paper's credibility. However, the central empirical claim hinges on V2, and the current definition of V2 has load-bearing flaws that undercut the 'meaningful morphology compatibility' conclusion. The V1 result is a system guarantee rather than a measured outcome. With a revised V2 that uses direction-aware scoring, a null baseline, and calibration against known perturbagens, the framework could become a useful template; in its current form the quantitative evidence does not support the abstract's headline claims.

major comments (5)
  1. [§3.7.2, Eqs. (5)–(6), Table A6] V2 ignores the direction codes that Table A6 explicitly defines. For example, in Table 1 the observed glcm_contrast_median ΔZ is −0.268, but senescence_like expects glcm_contrast(+); under Eq. (5) this feature counts as a hit anyway, contributing to the reported Hit@10=0.875. Thus the metric cannot penalize biologically contradictory outputs, despite §3.7.2 stating that V2 is 'designed to penalize biologically contradictory LLM outputs.' The abstract's 'meaningful morphology compatibility' is therefore not established by the score as implemented.
  2. [§4.3, Table A6, §3.7.2] No null baseline is reported for V2. Many proxy sets are broad (senescence_like includes 8 of the 9 reduced base features) and contain direction-code-0 entries, so a random or vacuous label can achieve moderate Hit@10. In addition, the headline dose-stratified numbers are the 'best mean Hit@10' (max over dose×label rows), which inflates the apparent dose-response; the global mean Hit@10 increases only from 0.2292 to 0.5787. Without a permutation/random-label baseline, the reported trend 0.236→0.889 cannot be interpreted as evidence that compatibility increases with perturbation strength.
  3. [§3.5.3, §3.7.1, §4.1] V1 is perfect by construction: the LLM runner rejects and retries any hypothesis with missing or invalid references, so invalid citations are filtered before scoring. The paper acknowledges this in §4.1, but the abstract's statement 'V1 detected no invalid evidence references' presents a system-design property as though it were an empirical audit outcome. This should be qualified both in the abstract and in §4.1, distinguishing enforced constraints from measured performance.
  4. [§4.5, Appendix A.1] The only external anchor for V2 is the correlation with the Step 4 drift summary: Pearson r=0.307 with 95% CI [−0.001, 0.566] and Spearman ρ=0.293 with CI [−0.021, 0.558], i.e., both intervals include zero. The coarse dose-level analysis (n=5) has very wide CIs. Since both quantities derive from the same morphology distribution, this does not constitute strong independent validation. Combined with the V2 issues above, this is too weak to support 'meaningful' compatibility; the abstract's 'positively associated' should be tempered or replaced with a confidence-aware statement.
  5. [§3.4.4, §3.7.2, Table A6 vs Table 1] The V2 computation is underspecified. The prompt's observation block contains features such as 'mean_intensity_median' and 'glcm_energy_median', while Table A6 defines proxy sets over reduced features such as 'mean_intensity' and 'glcm_energy'. The manuscript never states the mapping rule from the 36 well-level summary features to the reduced proxy vocabulary, nor the size of T (the representative payload in Table 1 has 8 features, yet Eq. (5) is called Hit@10). This is a reproducibility gap; the reported V2 values cannot be reconstructed from the text. Please specify the matching rule, the exact T construction, and the handling of summary statistics.
minor comments (4)
  1. [References] References [9] and [10] appear to be duplicates (Hernandez-Segura et al., 2018), as do [23] and [24] (Neurohr et al., 2019). Please deduplicate.
  2. [Table A6] The label 'other_uncertain' has an empty proxy set, so any hypothesis with this label necessarily receives Hit@10=0 and WP=0. The paper should describe how such hypotheses are treated in the per-dose aggregations, since their inclusion will deflate means and may affect the dose-response.
  3. [Eq. (5)] If |T| is not always 10, the name 'Hit@10' is misleading. Either fix T to exactly 10 features or rename the metric (e.g., Hit@k) and report the distribution of |T|.
  4. [§3.4.3, Table A2] Pathway references are listed as citable for hierarchical levels but not for week×dose prompts. Clarify whether path:<id> references are ever present in week×dose payloads or only in dose_time/global prompts, to avoid ambiguity in V1 coverage.

Circularity Check

1 steps flagged

V1 'no invalid references' is guaranteed by the reject/retry grounding filter; V2 is internal but not circular as a consistency metric.

specific steps
  1. self definitional [Section 3.5.3 (Grounding enforcement), Section 3.5.1 (Model and decoding), Section 4.1 (Grounding integrity), Eq. (4)]
    "Grounding enforcement. For week ×dose hypotheses, grounding validation requires that all quant_evidence_refs are of the form obs:<obs_id> and match the prompt’s observation ID, and that all retrieval_refs are of the form jump:<jump_id> or lit:<lit_id> and appear in the payload. ... The runner retries up to three times on schema or grounding failures. ... By construction, these V1 results establish provenance validity within the closed prompt evidence set; they do not by themselves show that a cited item semantically supports a claim."

    V1 is not an independent audit of the LLM's citations: the same subset/format predicate is the acceptance criterion used during generation, with failed outputs rejected and retried. Any output that survives to scoring satisfies V1 by construction, so the abstract's 'V1 detected no invalid evidence references' is an identity between the metric and the system's filter, not an empirical finding about citation behavior. The paper's own 'By construction' sentence makes this reduction explicit.

full rationale

The only load-bearing step that reduces to its own inputs is the V1 grounding result. Everything downstream (V2, drift correlation, hierarchical summaries) is either an internal consistency metric or is explicitly qualified. V2's proxy table is hand-specified from external morphological hallmarks, the LLM never sees the proxy table, and no target values are fitted to the V2 scores, so V2 is not circular in the derivation sense; however, the skeptic's objections (Eq. 5 ignores Table A6 direction codes, no null baseline, proxy set is internal) are correctness/calibration concerns that weaken the 'meaningful morphology compatibility' claim without making it circular. The drift comparison is a descriptive correlation between two functions of the same morphology deltas, and the paper repeatedly states it is not independent biological validation. No self-citation chain or imported uniqueness theorem is load-bearing. Score 5 reflects one headline auditing result that is true by construction while the central framework retains independent content.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 1 invented entities

The central validation rests on the hand-authored V2 proxy table, the reject/retry grounding loop that makes V1 trivially perfect, and retrieval relevance assumptions. No new physical forces or particles are postulated; the 'adaptive phenotype' is an LLM-generated interpretive label rather than an independently evidenced entity.

free parameters (6)
  • n_PC (PCA components for drift summary) = 12
    Number of whitened PCA components chosen to summarize morphology drift (Section 3.3.1); affects the external anchor used in the V2 consistency check.
  • Retrieval hyperparameters (k=100, retain=10, min gene neighbors=3, top pathways=12) = 100 / 10 / 3 / 12
    Hand-chosen settings shaping the evidence payload; different values could change neighbors and downstream hypotheses (Section 3.4, Table A2).
  • TF-IDF retrieval settings (max_features=250000, min_df=2, ngram 1-2) = 250000 / 2
    Used for Reactome pathway retrieval; hand-chosen (Section 3.4.3, Table A2).
  • LLM decoding settings (temperature=0.2, max tokens=8192, thinking=HIGH) = 0.2 / 8192 / HIGH
    Affects hypothesis diversity and retry behavior; not tuned against V2 (Section 3.5.1).
  • V2 proxy feature sets and direction codes per process label = Table A6 (e.g., senescence_like includes area+, perimeter+, glcm_contrast+, glcm_energy-, ...)
    Hand-specified literature-based mapping; the entire V2 metric depends on it, and direction codes are not enforced by Eq. 5-6.
  • Reduced 9-feature JUMP embedding (averaging over channels/scales) = 9 dimensions
    Mapping Cell Painting columns to the 9 base features is a modeling choice that affects neighbor retrieval (Section 3.4.1).
axioms (6)
  • domain assumption Cell Painting morphology features are interpretable proxies for biological processes (senescence, apoptosis, DDR, etc.) as encoded in Table A6.
    Invoked in Section 3.7.2 and Table A6; if false, V2 cannot measure biological compatibility.
  • domain assumption The weekly paired control baseline and within-week robust normalization remove technical variation without erasing treatment signal.
    Step 3 (Eq. 1-2); the treated-control deltas are the foundation for all downstream evidence and retrieval.
  • domain assumption JUMP reference profiles, Reactome pathways, and EuropePMC snippets retrieved for each week x dose payload are relevant to RPE-1 low-dose-rate radiation response.
    Section 3.4; retrieval relevance is not quantitatively validated against ground truth.
  • ad hoc to paper Strict grounding enforcement with reject/retry is the right way to achieve V1; the perfect V1 result is therefore by construction.
    Section 3.5.3 and Section 4.1 explicitly state V1 is by construction; this makes the 'perfect V1' headline a property of the pipeline rather than an empirical test.
  • ad hoc to paper The 11-label controlled vocabulary is sufficient to capture relevant biological processes for this dataset.
    Table A2/A4; coarse labels may collapse distinct mechanisms, as acknowledged in Section 5.
  • standard math PCA/whitening on control wells gives a valid low-dimensional summary of morphology drift.
    Standard linear method; the specific n_PC=12 choice is in Section 3.3.1.
invented entities (1)
  • Low-dose adaptive phenotype (metabolic reprogramming + proteostatic stress at 0.003-0.3 mGy/hr) no independent evidence
    purpose: Interpretive hypothesis explaining lower-dose morphology trajectories.
    Produced by the LLM/hierarchical summary (Section 4.6, Table A7); no orthogonal experimental validation, and the authors explicitly say results do not establish mechanisms.

pith-pipeline@v1.3.0-alltime-deepseek · 17813 in / 16586 out tokens · 131937 ms · 2026-08-01T21:15:10.083548+00:00 · methodology

0 comments
read the original abstract

High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology remains difficult, especially for weak, chronic perturbations such as low-dose-rate ionizing radiation. Large language models (LLMs) can synthesize heterogeneous evidence into biological narratives, yet their scientific use requires quantitative auditing. We present an evaluation-first, retrieval-augmented interpretation framework for longitudinal Cell Painting morphology, applied to a 9-week RPE-1 time course across five dose rates (0.003--6.0 mGy/hr). Week-matched treated-control morphology deltas are combined with retrieved perturbation neighbors, pathway context, and literature evidence through stable evidence identifiers, enabling an LLM to generate structured, evidence-linked hypotheses that are hierarchically summarized while preserving provenance. We introduce two quantitative auditing tests: V1 citation validity, which verifies that cited evidence identifiers exist in the prompt, and V2 proxy-based morphology compatibility, which evaluates consistency between predicted biological processes and the most altered morphology features. In our experiments, V1 detected no invalid evidence references, while V2 showed meaningful morphology compatibility that increased with perturbation strength and was positively associated with an independent morphology drift summary. The framework produces auditable, falsifiable biological hypotheses, including an adaptive phenotype involving metabolic reprogramming and proteostatic stress at lower dose rates (0.003--0.3 mGy/hr). Current limitations include proxy-based evaluation and the lack of ground-truth mechanism labels.

Figures

Figures reproduced from arXiv: 2607.19415 by Byung-Jun Yoon, Gilchan Park, Guang Zhao, Shinjae Yoo.

Figure 1
Figure 1. Figure 1: Architectural overview of the evaluation-first LLM-driven morphological reasoning and validation pipeline. The [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: V2 proxy alignment dose-response (Hit@10). [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Dose-time integrated phase timeline (hierarchical reasoning output). [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Temporal heatmap of LLM process_label assign￾ments. features. Such changes are consistent with stress-associated mor￾phological remodeling and hypertrophic phenotypes commonly ob￾served during cellular senescence and growth dysregulation [9, 24]. Importantly, prior work in image-based profiling has demon￾strated that morphological features extracted from Cell Painting assays can reliably capture biological… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 3 canonical work pages · 1 internal anchor

  1. [1]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623

  2. [2]

    Daniil A Boiko, Robert MacKnight, and Gabe Gomes. 2023. Emergent au- tonomous scientific research capabilities of large language models.arXiv preprint arXiv:2304.05332(2023)

  3. [3]

    Davis, Blake Borgeson, Carol Hartland, Maria Kost-Alimova, Sigrun M

    Mark-Anthony Bray, Shantanu Singh, Han Han, Chad T. Davis, Blake Borgeson, Carol Hartland, Maria Kost-Alimova, Sigrun M. Gustafsdottir, Christopher C. Gibson, and Anne E. Carpenter. 2016. Cell Painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes.Nature Protocols11, 9 (2016), 1757–1774. doi:10.1038/nprot...

  4. [4]

    Juan C Caicedo, Sam Cooper, Florian Heigwer, Scott Warchal, Peng Qiu, Csaba Molnar, Aliaksei S Vasilevich, Joseph D Barry, Harmanjit Singh Bansal, Oren Kraus, et al. 2017. Data-analysis strategies for image-based cell profiling.Nature methods14, 9 (2017), 849–863

  5. [5]

    Cimini, Amy Goodale, Lisa Miller, Maria Kost-Alimova, Nasim Jamali, John G

    Srinivas Niranj Chandrasekaran, Beth A. Cimini, Amy Goodale, Lisa Miller, Maria Kost-Alimova, Nasim Jamali, John G. Doench, Briana Fritchman, Adam Skepner, Michelle Melanson, Alexandr A. Kalinin, John Arevalo, Marzieh Haghighi, Juan C. Caicedo, Daniel Kuhn, Desiree Hernandez, James Berstler, Hamdah Shafqat- Abbasi, David E. Root, Susanne E. Swalley, Saksh...

  6. [6]

    Europe PMC Consortium. 2015. Europe PMC: a full-text literature database for the life sciences and platform for innovation.Nucleic acids research43, D1 (2015), D1042–D1048

  7. [7]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM64, 12 (2021), 86–92. doi:10.1145/3458723 Originally released as arXiv:1803.09010

  8. [8]

    Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2020. PubMedBERT: Domain-Specific Language Model Pretraining for Biomedical Text Mining. arXiv:2007.15779 [cs.CL] doi:10.48550/arXiv.2007.15779 TO BE VERIFIED: if citing a published version, replace arXiv with venue details

  9. [9]

    Alejandra Hernandez-Segura, Jamil Nehme, and Marco Demaria. 2018. Hallmarks of cellular senescence.Trends in cell biology28, 6 (2018), 436–453

  10. [10]

    Arturo Hernandez-Segura, Justine Nehme, and Marco Demaria. 2018. New hallmarks of cellular senescence.Trends in Cell Biology28, 6 (2018), 436–453

  11. [11]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Hao- tian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A Survey on Hallucination in Large Language Models: Prin- ciples, Taxonomy, Challenges, and Open Questions. arXiv:2311.05232 [cs.CL] doi:10.48550/arXiv.2311.05232 arXiv:2311.05232 is a hallucination su...

  12. [12]

    Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs.IEEE transactions on big data7, 3 (2019), 535–547

  13. [13]

    Guido Kroemer, Lorenzo Galluzzi, Peter Vandenabeele, John Abrams, Emad S Alnemri, Eric H Baehrecke, Mikhail V Blagosklonny, Wafik S El-Deiry, Pierre Gol- stein, Douglas R Green, et al. 2009. Classification of cell death: recommendations of the Nomenclature Committee on Cell Death 2009.Cell Death & Differentiation 16, 1 (2009), 3–11

  14. [14]

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chanho So, and Jaewoo Kang. 2020. BioBERT: a pre-trained biomedical language representation model for biomedical text mining.Bioinformatics36, 4 (2020), 1234–1240. doi:10.1093/bioinformatics/btz682

  15. [15]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. InAdvances in Neural Informa- tion Processing Systems 33 (NeurIPS 2020). https://proceedings.ne...

  16. [16]

    Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. InProceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers). 3214–3252

  17. [17]

    Renqian Luo, Litao Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. BioGPT: generative pre-trained transformer for biomedical text generation and mining. arXiv:2210.10341 [cs.CL] doi:10.48550/arXiv.2210.10341

  18. [18]

    Markus Marks, Uriah Israel, Rohit Dilip, Qilin Li, Changhua Yu, Emily Laubscher, Ahamed Iqbal, Elora Pradhan, Ada Ates, Martin Abt, et al . 2025. CellSAM: a foundation model for cell segmentation.Nature Methods(2025), 1–9

  19. [19]

    Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. InProceedings of the 58th annual meeting of the association for computational linguistics. 1906–1919

  20. [20]

    Marija Milacic, Deidre Beavers, Patrick Conley, Chuqiao Gong, Marc Gille- spie, Johannes Griss, Robin Haw, Bijay Jassal, Lisa Matthews, Bruce May, Robert Petryszak, Eliot Ragueneau, Karen Rothfels, Cristoffer Sevilla, Veron- ica Shamovsky, Ralf Stephan, Krishna Tiwari, Thawfeek Varusai, Joel Weiser, Adam Wright, Guanming Wu, Lincoln Stein, Henning Hermjak...

  21. [21]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. InProceedings of the Conference on Fairness, Ac- countability, and Transparency (FAT* ’19). 220–229. doi:10.1145/3287560.3287596

  22. [22]

    2006.Health Risks from Exposure to Low Levels of Ionizing Radiation: BEIR VII Phase 2

    National Research Council. 2006.Health Risks from Exposure to Low Levels of Ionizing Radiation: BEIR VII Phase 2. The National Academies Press, Washington, DC. doi:10.17226/11340

  23. [23]

    Gabriel E Neurohr, Rachel L Terry, Jette Lengefeld, Megan Bonney, Gregory P Brittingham, Fabrice Moretto, Teemu P Miettinen, Laura P Vaites, Lidia M Soares, Joao A Paulo, et al. 2019. Excessive cell growth causes cytoplasm dilution and contributes to senescence.Cell176, 5 (2019), 1083–1097

  24. [24]

    Gabriel E Neurohr, Rachel L Terry, Jette Lengefeld, Megan Bonney, Gregory P Brittingham, Fabien Moretto, Teemu P Miettinen, Laura Pontano Vaites, Luis M Soares, Joao A Paulo, et al. 2019. Excessive cell growth causes cytoplasm dilution and contributes to senescence.cell176, 5 (2019), 1083–1097

  25. [25]

    Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al . 2011. Scikit-learn: Machine learning in Python.the Journal of machine Learning research12 (2011), 2825–2830

  26. [26]

    Mohammad Hasan Rohban, Shantanu Singh, Xiaoyun Wu, Julia B Berthet, Mark- Anthony Bray, Matthew D Shair, Lee L Rubin, and Anne E Carpenter. 2017. Systematic morphological profiling of human gene and allele function via Cell Painting.eLife6 (2017), e24060

  27. [27]

    Mohammad Hossein Rohban, Shantanu Singh, Xiaoyun Wu, Julia B Berthet, Mark-Anthony Bray, Yashaswi Shrestha, Xaralabos Varelas, Jesse S Boehm, and Anne E Carpenter. 2017. Systematic morphological profiling of human gene and allele function via Cell Painting.Elife6 (2017), e24060

  28. [28]

    Maria Janina Sarol, Shufan Ming, Shruthan Radhakrishna, Jodi Schneider, and Halil Kilicoglu. 2024. Assessing citation integrity in biomedical publications: corpus annotation and NLP models.Bioinformatics40, 7 (2024), btae420. doi:10. 1093/bioinformatics/btae420 Published 2024-06-26

  29. [29]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, et al . 2023. Large language models encode clinical knowledge.Nature620, 7972 (2023), 172–180

  30. [30]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805(2023)

  31. [31]

    United Nations Scientific Committee on the Effects of Atomic Radiation. 2022. Sources, Effects and Risks of Ionizing Radiation: UNSCEAR 2020/2021 Report to the General Assembly, with Scientific Annexes, Volume III – Scientific Annex C: Biological mechanisms relevant for the inference of cancer risks from low-dose and low-dose-rate radiation. Technical Rep...

  32. [32]

    David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7534–7550. doi:10.18653/v1/2020.emnlp-main.609 Auditing Retrieval-Augmented LLM Hypotheses for ...