REVIEW 4 major objections 4 minor 98 references
Aligning LMs with UMLS subgraphs boosts biomedical QA and linking.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-08-04 21:55 UTC pith:SAMYO6C5
load-bearing objection BALI is a solid pre-training method for biomedical LMs, with real QA gains; the entity-linking improvements are partly circular and need an MLM-only control. the 4 major comments →
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
BALI's central claim is that explicit cross-modal alignment between a language model's entity representations and a graph neural network's concept representations transfers usable biomedical knowledge into the LM. For each mention in a masked sentence, the mention's pooled token embeddings serve as the textual anchor, and a GAT encoder produces a structural anchor from a 1-hop UMLS subgraph seeded with concept-name embeddings. An InfoNCE loss pulls paired anchors together while MLM preserves language ability. After this short pre-training, the graph encoder is discarded, yet the LM alone shows mean accuracy gains of 2.1, 1.7, and 6.2 points for PubMedBERT on PubMedQA, MedQA, and BioASQ, larg
What carries the argument
Cross-modal anchoring: a single biomedical concept is represented in two complementary modalities — a textual representation (mean-pooled token embeddings of its mention in context) and a structural representation (a GAT-encoded local UMLS subgraph, with initial node features obtained by encoding randomly sampled concept names with the LM). The InfoNCE contrastive loss over paired text-graph anchors is the alignment mechanism, MLM is retained to preserve language ability, and the graph encoder is discarded after pre-training.
Load-bearing premise
The alignment signal is only as good as the automatic linking that matches mentions in the training sentences to UMLS concepts; if many links are wrong, the contrastive objective pushes text embeddings toward the wrong graph nodes, and the reported gains may come mostly from continued masked-language training instead.
What would settle it
Pre-train BALI twice with identical data and compute, once with true mention-to-UMLS links and once with randomly shuffled links; if the shuffled model keeps the QA and entity-linking gains, then the knowledge-graph alignment itself is not the driver and the effect comes from the extra pre-training exposure or MLM.
If this is right
- Knowledge can be injected during pre-training alone, so downstream tasks need no retrieved subgraphs or entity linking at inference time.
- A small, balanced alignment corpus (1.67M sentences, ~600K concepts, 65K steps, roughly 9 GPU-hours for base models) is sufficient for consistent QA improvements.
- Zero-shot entity linking improves sharply for general biomedical LMs, with Accuracy@1 rising by about 13 points for PubMedBERT and 24 points for BioLinkBERT-base on average across five corpora.
- A BALI-adapted BioLinkBERT-base can match or slightly beat SapBERT, a task-specific model pre-trained on the full UMLS synonym space, on supervised entity linking.
- GAT-based subgraph aggregation outperforms mean-pooling graph encoders, while larger LMs can benefit more from a single-encoder linearized-graph variant.
Where Pith is reading between the lines
- The paper's own ablation shows that removing the alignment loss but keeping MLM already yields 63.78 on PubMedQA versus the 63.1 raw PubMedBERT baseline, so part of the gain is plausibly from continued masked-language pre-training rather than graph alignment; no fully matched equal-compute control is reported.
- Because the training data relies on automatic entity recognition and normalization, incorrect mention-to-concept links would push text embeddings toward wrong graph nodes; a gold-annotated or confidence-filtered training set would likely sharpen the measured alignment signal.
- The same anchor-based recipe should transfer to other text-attributed knowledge graphs, with the effective graph encoder choice depending on model capacity — small encoders appear to need the external GAT, larger ones can use linearized subgraphs.
- A direct negative-control experiment, pairing sentences with randomly shuffled subgraphs instead of their true linked subgraphs, would isolate whether the contrastive alignment itself or the extra pre-training data drives the reported gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes BALI, a pre-training method that augments biomedical language models (PubMedBERT, BioLinkBERT base/large) by aligning pooled entity-mention representations with UMLS knowledge-graph subgraph representations. A GAT encoder encodes local 1-hop subgraphs, whose initial node features are LM-encoded concept names; an InfoNCE contrastive loss pulls mention and graph representations of the same concept together, while MLM is retained as a joint objective. After pre-training on 1.67M PubMed sentences with BERN2-based entity linking, the GNN is discarded, and the resulting LM is evaluated on biomedical question answering (PubMedQA, MedQA, BioASQ), entity linking (zero-shot and supervised), and relation extraction (ChemProt, DDI, GAD). The paper reports consistent QA gains over base models, large zero-shot entity-linking gains for general biomedical LMs, small relation-extraction gains, and ablations showing that both objectives and attention-based graph aggregation matter.
Significance. If the claims hold, BALI would be a valuable and relatively cheap recipe for infusing structured biomedical knowledge into lightweight encoders without any inference-time KG access. The paper has concrete strengths: it releases code and pre-trained models, reports hyperparameters in Table 1, averages repeated fine-tuning runs for PubMedQA and BioASQ, and includes a useful ablation suite (Table 6) covering loss removal, alternative alignment losses, aggregation methods, and GNN depth. The core idea of explicit cross-modal alignment through entity anchors is clearly presented and is distinct from prior interaction-token or retrieval-augmented approaches. However, the current evidence for the entity-representation claim is weakened by the absence of an MLM-only control in the entity-linking experiments and by the close similarity between the zero-shot entity-linking protocol and the training objective. The QA results are more independent and, with the existing MLM-only ablation, already show that the alignment term contributes beyond continued pre-training, which is a useful partial control.
major comments (4)
- [§4.2, §5.3.2, Table 3] The zero-shot entity-linking evaluation largely re-measures the training objective. L_align (Eq. 5) maximizes cosine similarity between a pooled mention representation and a GNN subgraph representation whose initial node features are LM-encoded concept names (§4.1). The zero-shot retrieval protocol (§5.3.2) retrieves by cosine similarity between mention and concept-name representations. This is close to the positive-pair construction in training. More importantly, Table 3 reports no MLM-only control for entity linking, even though Table 6 shows that MLM-only accounts for a substantial portion of the QA gains (PubMedQA 63.1→63.78; BioASQ 67.8→70.58). Without an equivalent EL control or a disjoint-concept test, the headline EL gains (e.g., PubMedBERT +13.1 Accuracy@1) cannot be cleanly attributed to the KG-alignment term rather than to continued MLM on the same entity-linked corpus.
- [§5.3.1, Table 2] The text claims that BioLinkBERT_large 'after BALI pretraining performs on par or better than the task-specific QA-GNN and GreaseLM methods.' This is contradicted by Table 2: BioLinkBERT_large+BALI(GNN) scores 68.7 on PubMedQA and 45.0 on MedQA, while QA-GNN scores 72.1/45.0 and GreaseLM 72.4/45.1. Even the linear-graph variant (70.9 PubMedQA) is below QA-GNN and GreaseLM. The comparison should be corrected, or the claim narrowed to specific datasets where it actually holds (e.g., MedQA).
- [§5.1 / Table 2] MedQA results are reported as single means with no standard deviations or error bars, whereas PubMedQA and BioASQ have ± values averaged over 10 runs. Since the reported MedQA gains are small (e.g., PubMedBERT 38.1→39.8; BioLinkBERT_large 44.6→45.0), it is impossible to assess whether these differences are significant. Please report variance or at least the number of seeds used for MedQA.
- [§5, 'Pretraining Data'] The pre-training corpus is built with BERN2 for named entity recognition and UMLS normalization, but the paper reports no quality check on this linking. If a substantial fraction of mentions are mapped to incorrect UMLS concepts, the contrastive loss will pull text representations toward wrong graph nodes, and the observed gains would reflect a different effect. The manuscript should at least report a sample-based precision estimate for the BERN2 linking on the pre-training data, or provide a robustness analysis (e.g., training on a filtered subset with high-confidence links).
minor comments (4)
- [§5.1] The evaluation-task list has a duplicate: item (iii) is listed as 'BC5CDR-D' but should presumably be the BC5CDR-Chemical corpus, matching Table 3's 'BC5CDR-C' column.
- [§5.5.3 / Table 6] The ablation text says 'larger (5 layers)' but Table 6 reports L=7 for the larger GNN. The text and table should be harmonized.
- [Abstract / §1 / §5 / Conclusion] The pre-training corpus size is given inconsistently as '1.5M sentences' (Introduction), '1.67M sentences' (Pretraining Data), and '1.7M sentences' (Conclusion). Please use one consistent figure.
- [Table 3] The underline convention ('best of two scores') is defined only for original vs. BALI-pretrained models. For the SapBERT and GEBERT rows, the notational significance of underlining is less clear; please clarify or add bolding for the overall best per column.
Circularity Check
Zero-shot entity-linking evaluation largely re-measures the BALI alignment objective; QA/RE provide independent support, so circularity is partial.
specific steps
-
fitted input called prediction
[Section 4.2 (L_align) and Section 5.3.2 (zero-shot retrieval protocol)]
"L_align = − 1/B Σ_i log exp(cos(ē_i, ḡ_i)/τ) / Σ_j exp(cos(ē_j, ḡ_j)/τ) ... As an initial representation for a node u, a random concept name s_u ∈ S_u is sampled and encoded with a textual encoder: ḡ_u^(0) = LM(s_u). ... zero-shot similarity-based retrieval approach over pooled mention and concept name representations [74]."
The zero-shot EL protocol retrieves concepts by cosine similarity between a pooled mention representation and a concept-name representation. The BALI alignment loss L_align directly maximizes cosine similarity between the same pooled mention representation ē_v and a graph representation ḡ_v whose initial node feature is the LM-encoded concept name (ḡ_u^(0)=LM(s_u)). After training, the GNN output is a function of these concept-name embeddings, so the mention representations have been pulled toward concept-name-derived vectors by the training objective itself. Reporting large zero-shot EL gains as evidence that 'external KG structure was internalized' is therefore partly circular: the evaluation metric is essentially the training similarity, not an independent probe of graph structure. The
full rationale
The paper's central claims are supported by a mix of independent and partly circular evidence. The QA evaluations (PubMedQA, MedQA, BioASQ) and relation extraction (ChemProt, DDI, GAD) are external downstream tasks not defined by the alignment objective; the ablation in Table 6 shows that removing L_align (MLM-only) yields 63.78 on PubMedQA and 70.58 on BioASQ vs 63.1/67.8 for the base model, so continued LM pre-training accounts for part of the gain, but the full BALI model (65.2/74) still exceeds the MLM-only control, giving the alignment term some independent content. The zero-shot entity linking results in Table 3 are the main evidence for 'quality of entity representations,' and these are close to the training objective: the model is trained with an InfoNCE loss on cos(mention, graph-concept) and tested with cosine retrieval against concept names, where the graph representation is initialized from LM-encoded concept names. This is a partial reduction by construction, but not a full one because the GNN aggregates neighboring concept-name embeddings and the evaluation corpora are not identical to the pre-training sentences. The absence of an MLM-only control for entity linking is a gap but not itself a circular step. No load-bearing self-citation chain was found: references to the authors' prior GEBERT work ([58]) support GAT choice and MS-loss hyperparameters but are not the basis for the main empirical claim. No uniqueness theorem or ansatz smuggled via citation. Overall, circularity is partial and limited to the entity-linking evaluation; the paper is otherwise self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
free parameters (4)
- InfoNCE temperature tau =
not stated
- GNN hidden size, layers, heads =
768, 5 layers, 2 heads
- Max neighbors per node =
3
- LM and non-LM learning rates =
2e-5 and 1e-4
axioms (3)
- domain assumption UMLS concept names and relations are a reliable source of biomedical knowledge.
- domain assumption BERN2 entity recognition and normalization on PubMed abstracts is accurate enough for the contrastive pairs.
- ad hoc to paper Aligning pooled entity representations improves the LM more broadly, not just the entity pool.
Cite this review
Pith. "Pith review of BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment." pith.science (2026). https://pith.science/paper/SAMYO6C5
@misc{pith2026250907588,
author = {Pith},
title = {Pith review of: BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/SAMYO6C5}},
note = {Machine review of arXiv:2509.07588}
}
read the original abstract
In recent years, there has been substantial progress in using pretrained Language Models (LMs) on a range of tasks aimed at improving the understanding of biomedical texts. Nonetheless, existing biomedical LLMs show limited comprehension of complex, domain-specific concept structures and the factual information encoded in biomedical Knowledge Graphs (KGs). In this work, we propose BALI (Biomedical Knowledge Graph and Language Model Alignment), a novel joint LM and KG pre-training method that augments an LM with external knowledge by the simultaneous learning of a dedicated KG encoder and aligning the representations of both the LM and the graph. For a given textual sequence, we link biomedical concept mentions to the Unified Medical Language System (UMLS) KG and utilize local KG subgraphs as cross-modal positive samples for these mentions. Our empirical findings indicate that implementing our method on several leading biomedical LMs, such as PubMedBERT and BioLinkBERT, improves their performance on a range of language understanding tasks and the quality of entity representations, even with minimal pre-training on a small alignment dataset sourced from PubMed scientific abstracts.
Figures
Reference graph
Works this paper leans on
-
[1]
Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019. Publicly Available Clinical BERT Embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Work- shop, Anna Rumshisky, Kirk Roberts, Steven Bethard, and Tristan Naumann (Eds.). Association for Computational Linguistics, Minne...
-
[2]
Jinheon Baek, Alham Fikri Aji, and Amir Saffari. 2023. Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answer- ing. In Proceedings of the 1st Workshop on Natural Language Reasoning and Struc- tured Explanations (NLRSE). 78–106
2023
-
[3]
Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, and Yulia Tsvetkov. 2024. Kgquiz: Evaluating the generalization of encoded knowledge in large language models. In Proceedings of the ACM on Web Conference 2024. 2226–2237
2024
-
[4]
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text. InProceedings of the 2019 Conference on Empirical Meth- ods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association...
-
[5]
Olivier Bodenreider. 2004. The Unified Medical Language System (UMLS): inte- grating biomedical terminology. Nucleic Acids Research 32, Database-Issue (2004), 267–270. doi:10.1093/NAR/GKH061
-
[6]
Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, Un...
2013
-
[7]
Àlex Bravo, Janet Piñero, Núria Queralt-Rosinach, Michael Rautschka, and Laura I Furlong. 2015. Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research. BMC bioinfor- matics 16 (2015), 1–17
2015
-
[8]
Shaked Brody, Uri Alon, and Eran Yahav. 2022. How Attentive are Graph At- tention Networks?. In The Tenth International Conference on Learning Repre- sentations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net. https: //openreview.net/forum?id=F72ximsx7C1
2022
-
[9]
David Chang, Ivana Balazevic, Carl Allen, Daniel Chawla, Cynthia Brandt, and Richard Andrew Taylor. 2020. Benchmark and Best Practices for Biomedical Knowledge Graph Embeddings. InProceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing, BioNLP 2020, Online, July 9, 2020 , Dina Demner- Fushman, Kevin Bretonnel Cohen, Sophia Ananiadou, ...
-
[10]
Qijie Chen, Haotong Sun, Haoyang Liu, Yinghui Jiang, Ting Ran, Xurui Jin, Xianglu Xiao, Zhimin Lin, Hongming Chen, and Zhangmin Niu. 2023. An extensive benchmark study on biomedical text generation and min- ing with ChatGPT. Bioinformatics 39, 9 (09 2023), btad557. doi:10.1093/ bioinformatics/btad557 arXiv:https://academic.oup.com/bioinformatics/article- ...
2023
-
[11]
David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Ko- zlova, Nikita Semenov, and Alexander Panchenko. 2021. Text Detoxification using Large Pre-trained Neural Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds...
-
[12]
Allan Peter Davis, Thomas C Wiegers, Michael C Rosenstein, and Carolyn J Mattingly. 2012. MEDIC: a practical disease vocabulary used at the Comparative Toxicogenomics Database. Database 2012 (2012), bar065
2012
-
[13]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, ...
2019
-
[14]
Rezarta Islamaj Dogan, Robert Leaman, and Zhiyong Lu. 2014. NCBI disease corpus: A resource for disease name recognition and concept normalization. J. Biomed. Informatics 47 (2014), 1–10. doi:10.1016/J.JBI.2013.12.006
-
[15]
Matthias Fey and Jan E. Lenssen. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds
2019
-
[16]
Schoenholz, Patrick F
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017. Neural Message Passing for Quantum Chemistry. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research, Vol. 70) , Doina Precup and Yee Whye Teh (Eds.)...
2017
-
[17]
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2022. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans. Comput. Heal. 3, 1 (2022), 2:1–2:23. doi:10.1145/3458754
doi:10.1145/3458754 2022
-
[18]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning. PMLR, 3929–3938
2020
-
[19]
Hamilton, Zhitao Ying, and Jure Leskovec
William L. Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive Represen- tation Learning on Large Graphs. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S....
2017
-
[20]
Bin He, Di Zhou, Jinghui Xiao, Xin Jiang, Qun Liu, Nicholas Jing Yuan, and Tong Xu. 2020. Integrating Graph Contextualized Knowledge into Pre-trained Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020), Trevor Cohn, Yulan He, and Yang Liu (Eds.). As...
-
[21]
María Herrero-Zazo, Isabel Segura-Bedmar, Paloma Martínez, and Thierry De- clerck. 2013. The DDI corpus: An annotated corpus with pharmacological sub- stances and drug–drug interactions. Journal of biomedical informatics 46, 5 (2013), 914–920
2013
-
[22]
Chun-Yu Hsueh, Yu Zhang, Yu-Wei Lu, Jen-Chieh Han, Wilailack Meesawad, and Richard Tzong-Han Tsai. 2023. NCU-IISR: Prompt Engineering on GPT-4 to Stove Biological Problems in BioASQ 11b Phase B. In Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023), Thessaloniki, Greece, September 18th to 21st, 2023 (CEUR Workshop Proceedings, Vo...
2023
-
[23]
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences 11, 14 (2021). doi:10.3390/app11146421
-
[24]
Cohen, and Xinghua Lu
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu
-
[25]
Minki Kang, Jinheon Baek, and Sung Ju Hwang. 2022. KALA: Knowledge- Augmented Language Model Adaptation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguis- tics: Human Language Technologies, NAACL 2022, Seattle, W A, United States, July 10-15, 2022 , Marine Carpuat, Marie-Catherine de Marneffe...
2022
-
[26]
Seyed Mehran Kazemi and David Poole. 2018. SimplE Embedding for Link Prediction in Knowledge Graphs. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada , Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicolò Cesa-Bian...
2018
-
[27]
Pei Ke, Haozhe Ji, Yu Ran, Xin Cui, Liwei Wang, Linfeng Song, Xiaoyan Zhu, and Minlie Huang. 2021. JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 (Findings of ACL, Vol. ACL/IJCNLP 2021), Chengqing Zong...
-
[28]
Jing Yu Koh, Daniel Fried, and Russ Salakhutdinov. 2023. Generating Im- ages with Multimodal Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz H...
2023
-
[29]
Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. 2023. Grounding Language Models to Images for Multimodal Inputs and Outputs. InInternational Conference BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment SIGIR ’25, July 13–18, 2025, Padua, Italy on Machine Learning, ICML 2023, 23-29 July 2023, Hono...
work page 2023
-
[30]
Martin Krallinger, Obdulia Rabal, Saber A Akhondi, Martın Pérez Pérez, Jesús Santamaría, Gael Pérez Rodríguez, Georgios Tsatsaronis, Ander Intxaurrondo, José Antonio López, Umesh Nandal, et al . 2017. Overview of the BioCreative VI chemical-protein interaction Track. In Proceedings of the sixth BioCreative challenge evaluation workshop, Vol. 1. 141–146
work page 2017
-
[31]
Andrey Kutuzov, Mohammad Dorgham, Oleksiy Oliynyk, Chris Biemann, and Alexander Panchenko. 2019. Learning Graph Embeddings from WordNet-based Similarity Measures. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019) , Rada Mihalcea, Ekaterina Shutova, Lun- Wei Ku, Kilian Evang, and Soujanya Poria (Eds.). Associa...
-
[32]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems 33 (2020), 9459–9474
2020
-
[33]
Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J
Jiao Li, Yueping Sun, Robin J. Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J. Mattingly, Thomas C. Wiegers, and Zhiyong Lu. 2016. BioCreative V CDR task corpus: a resource for chemical disease relation extraction. Database J. Biol. Databases Curation 2016 (2016). doi:10.1093/DATABASE/BAW068
- [34]
-
[35]
Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Col- lier. 2021. Self-Alignment Pretraining for Biomedical Entity Representations. In Proceedings of the 2021 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021 , Kristina Tout...
2021
-
[36]
Fangyu Liu, Ivan Vulic, Anna Korhonen, and Nigel Collier. 2021. Learning Domain-Specialised Representations for Cross-Lingual Biomedical Entity Link- ing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Pro- cessing, ACL/IJCNLP 2021, (Volume 2: Short...
2021
-
[37]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Vi- sual Instruction Tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey L...
work page 2023
-
[38]
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2020. K-BERT: Enabling Language Representation with Knowledge Graph. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educationa...
-
[39]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)
Pith/arXiv arXiv 2019
-
[40]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Or- leans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/forum? id=Bkg6RiCqY7
work page 2019
-
[41]
Donna Maglott, Jim Ostell, Kim D Pruitt, and Tatiana Tatusova. 2007. Entrez Gene: gene-centered information at NCBI. Nucleic acids research 35, suppl_1 (2007), D26–D31
work page 2007
-
[42]
Aidan Mannion, Didier Schwab, and Lorraine Goeuriot. 2023. UMLS-KGI- BERT: Data-Centric Knowledge Integration in Transformers for Biomedical Entity Recognition. In Proceedings of the 5th Clinical Natural Language Processing Workshop, ClinicalNLP@ACL 2023, Toronto, Canada, July 14, 2023, Tristan Nau- mann, Asma Ben Abacha, Steven Bethard, Kirk Roberts, and...
doi:10.18653/v1/2023 2023
-
[43]
Zaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, and Nigel Collier. 2022. Rewire-then-Probe: A Contrastive Recipe for Probing Biomed- ical Knowledge of Pre-trained Language Models. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, ...
work page 2022
-
[44]
Chen, and Alexan- der Wong
George Michalopoulos, Yuanxin Wang, Hussam Kaka, Helen H. Chen, and Alexan- der Wong. 2021. UmlsBERT: Clinical Domain Knowledge Augmentation of Con- textual Embeddings Using the Unified Medical Language System Metathesaurus. In Proceedings of the 2021 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Langua...
2021
-
[45]
Fedor Moiseev, Zhe Dong, Enrique Alfonseca, and Martin Jaggi. 2022. SKILL: Structured Knowledge Infusion for Large Language Models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, W A, United States, July 10-15, 2022 , Marine Carpuat, Ma...
-
[46]
Alexander A Morgan, Zhiyong Lu, Xinglong Wang, Aaron M Cohen, Juliane Fluck, Patrick Ruch, Anna Divoli, Katrin Fundel, Robert Leaman, Jörg Hakenberg, et al. 2008. Overview of BioCreative II gene normalization. Genome biology 9 (2008), 1–19
work page 2008
-
[47]
Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Salvador Lima- López, Eulàlia Farré-Maduell, Luis Gascó, Martin Krallinger, and Georgios Paliouras. 2023. Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering. In Experi- mental IR Meets Multilinguality, Multimodality, and Int...
work page 2023
-
[48]
Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, and Georgios Paliouras. 2023. Overview of BioASQ Tasks 11b and Synergy11 in CLEF2023. In Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2023), Thessaloniki, Greece, September 18th to 21st, 2023 (CEUR Workshop Proceedings, Vol. 3497), Mohammad Aliannejadi, Guglielmo Faggi...
work page 2023
-
[49]
John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. 2008. Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for? Queue 6, 2 (2008), 40–53
work page 2008
-
[50]
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. 2023. Capabilities of GPT-4 on Medical Challenge Problems. CoRR abs/2303.13375 (2023). doi:10.48550/ARXIV.2303.13375 arXiv:2303.13375
-
[51]
OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023). doi:10. 48550/ARXIV.2303.08774 arXiv:2303.08774
-
[52]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Des- maison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, Hi...
work page 2019
-
[53]
Peters, Mark Neumann, Robert L
Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A. Smith. 2019. Knowledge Enhanced Contextual Word Representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Ho...
2019
-
[54]
Minh C. Phan, Aixin Sun, and Yi Tay. 2019. Robust Representation Learning of Biomedical Names. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics,...
-
[55]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, Virtual Event / Atlanta, Georgia, USA, November 9-19, 2020, Christine Cuicchi, Irene Qualters...
work page 2020
-
[56]
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deep- Speed: System Optimizations Enable Training Deep Learning Models with Over SIGIR ’25, July 13–18, 2025, Padua, Italy Andrey Sakhovskiy and Elena Tutubalina 100 Billion Parameters. In KDD ’20: The 26th ACM SIGKDD Conference on Knowl- edge Discovery and Data Mining, Virtual Event,...
work page 2020
-
[57]
Knowledge-Aware Language Model Pretraining
Corby Rosset, Chenyan Xiong, Minh Phan, Xia Song, Paul N. Bennett, and Saurabh Tiwary. 2020. Knowledge-Aware Language Model Pretraining. CoRR abs/2007.00655 (2020). arXiv:2007.00655 https://arxiv.org/abs/2007.00655
work page internal anchor Pith review Pith/arXiv arXiv 2020
-
[59]
doi:10.1109/SC41405.2020.00024
Pith/arXiv arXiv 2020
-
[60]
Mikhail Salnikov, Hai Le, Prateek Rajput, Irina Nikishina, Pavel Braslavski, Valentin Malykh, and Alexander Panchenko. 2023. Large Language Models Meet Knowledge Graphs to Answer Factoid Questions. In Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation, PACLIC 2023, The Hong Kong Polytechnic University, Hong Kong, SAR,...
work page 2023
-
[61]
Mohammad, Goran Nenadic, and Graciela Gonzalez-Hernandez
Abeed Sarker, Maksim Belousov, Jasper Friedrichs, Kai Hakala, Svetlana Kir- itchenko, Farrokh Mehryary, Sifei Han, Tung Tran, Anthony Rios, Ramakanth Kavuluru, Berry de Bruijn, Filip Ginter, Debanjan Mahata, Saif M. Mohammad, Goran Nenadic, and Graciela Gonzalez-Hernandez. 2018. Data and systems for medication-related text classification and concept norma...
work page 2018
-
[62]
Matthias Schildwächter, Alexander Bondarenko, Julian Zenker, Matthias Hagen, Chris Biemann, and Alexander Panchenko. 2019. Answering Comparative Ques- tions: Better than Ten-Blue-Links?. InProceedings of the 2019 Conference on Human Information Interaction and Retrieval, CHIIR 2019, Glasgow, Scotland, UK, March 10-14, 2019, Leif Azzopardi, Martin Halvey, ...
-
[63]
Özge Sevgili, Alexander Panchenko, and Chris Biemann. 2019. Improving Neural Entity Disambiguation with Graph Embeddings. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Work- shop, Fernando Alva-Manchego, Eunsol Choi, and Daniel Khashabi (Eds.). Associa- tion for Computational Linguistics, Flore...
doi:10.18653/v1/p19- 2019
-
[64]
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. REPLUG: Retrieval-Augmented Black-Box Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers) . 8364–8377
work page 2024
-
[65]
Andrey Sakhovskiy, Natalia Semenova, Artur Kadurin, and Elena Tutubalina
-
[66]
Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Ji- axiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. 2021. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137 (2021)
Pith/arXiv arXiv 2021
-
[67]
Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowl- edge Graph Embedding by Relational Rotation in Complex Space. In 7th Interna- tional Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net. https://openreview.net/forum?id=HkgEQnRqYQ
work page 2019
-
[68]
Mujeen Sung, Hwisang Jeon, Jinhyuk Lee, and Jaewoo Kang. 2020. Biomedical Entity Representations with Synonym Marginalization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). Association for Computational...
doi:10.18653/v1/ 2020
-
[69]
Mujeen Sung, Minbyul Jeong, Yonghwa Choi, Donghyeon Kim, Jin- hyuk Lee, and Jaewoo Kang. 2022. BERN2: an advanced neural biomedical named entity recognition and normalization tool. Bioin- formatics 38, 20 (09 2022), 4837–4839. doi:10.1093/bioinformatics/ btac598 arXiv:https://academic.oup.com/bioinformatics/article- pdf/38/20/4837/46535173/btac598.pdf
-
[70]
Yi, Minji Jeon, Sungdong Kim, and Jae- woo Kang
Mujeen Sung, Jinhyuk Lee, Sean S. Yi, Minji Jeon, Sungdong Kim, and Jae- woo Kang. 2021. Can Language Models be Biomedical Knowledge Bases?. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Marie-Francine Moens, Xuanjing Huang, Lucia ...
-
[71]
Yijun Tian, Huan Song, Zichen Wang, Haozhu Wang, Ziqing Hu, Fang Wang, Nitesh V Chawla, and Panpan Xu. 2024. Graph neural prompting with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19080–19088
work page 2024
-
[72]
Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang, and Zheng Zhang. 2020. CoLAKE: Contextualized Language and Knowledge Embedding. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020 , Do- nia Scott, Núria Bel, and Chengqing Zong (Eds.). Int...
-
[73]
Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. 2016. Complex Embeddings for Simple Link Prediction. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016 (JMLR Workshop and Conference Proceedings, Vol. 48), Maria-Florina Balcan and Kilian Q. We...
2016
-
[74]
Elena Tutubalina, Artur Kadurin, and Zulfat Miftahutdinov. 2020. Fair Evaluation in Concept Normalization: a Large-scale Comparative Analysis for BERT-based Models. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020 , Do- nia Scott, Núria Bel, and Chengqing Zong (Ed...
-
[75]
Elena Tutubalina, Artur Kadurin, and Zulfat Miftahutdinov. 2020. Fair evaluation in concept normalization: a large-scale comparative analysis for BERT-based mod- els. In Proceedings of the 28th International conference on computational linguistics . 6710–6716
work page 2020
-
[76]
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. ArXiv abs/1807.03748 (2018). https://api. semanticscholar.org/CorpusID:49670925
Pith/arXiv arXiv 2018
-
[77]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: An- nual Conference on Neural Information Processing Systems 2017, December 4- 9, 2017, Long Beach, CA, USA , Isabelle Guyon, Ulrike von Luxb...
work page 2017
-
[78]
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph Attention Networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net. https://openreview. net/forum?id=rJXMpikCZ
work page 2018
-
[79]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. LLaMA: Open and Efficient Foundation Language Models. CoRR abs/2302.13971 (2023). doi:10.48550/ARXIV.2302.13971 arXiv...
-
[80]
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. KEPLER: A Unified Model for Knowledge Embed- ding and Pre-trained Language Representation. Trans. Assoc. Comput. Linguistics 9 (2021), 176–194. doi:10.1162/TACL_A_00360
-
[81]
Xun Wang, Xintong Han, Weilin Huang, Dengke Dong, and Matthew R. Scott
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.