REVIEW 5 major objections 6 minor 81 references
K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor
T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read K-COMP claims that a compressor which generates short definitions of a question's medical terms before summarizing the retrieved passages makes retrieval-augmented answering more accurate than uncompressed retrieval or earlier compressors.
desk verdict A capable, well-benchmarked medical RAG compressor whose full system works, but the claimed knowledge-injection mechanism rests on an ablation that mostly shows the appended entity glosses doing the work, not the training objective. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the causal knowledge-injection objective used to train the compressor. The model's input is the question with its medical entities masked by a special token, concatenated with five retrieved passages, written $q_m \oplus P$; the training target is the entities, their one-line descriptions (each ended by an end-of-description token), and then the summary, generated in that order. Because the retrieved passages sit to the right of the masked spans, predicting the masks forces the model to use posterior context, the causal-masking idea taken from CM3, while the final summary is generated autoregressively with attention to the entities and descriptions it just produced. The descriptions themselves are harvested cheaply from the retrieval corpus: the first sentence of each MedCorp article is treated as a short description of its title, matched to the entities that ScispaCy detects in the question, so no manual annotation of prior knowledge is needed. The ablation of removing the descriptions from the reader prompt shows this machinery works through the injected knowledge rather than through better summarization.
What would settle it
Replace the generated entity descriptions with random first sentences from unrelated MedCorp articles, keeping everything else identical, and measure reader accuracy. If scores stay at the full K-COMP level, the gain comes from prompt structure rather than injected knowledge; if scores fall back to the no-description ablation level, the factual content of the descriptions carries the effect. This directly tests the load-bearing assumption because the paper's own Table 2 already shows accuracy falling when the descriptions are removed entirely.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that causal-masking training can turn a 2-billion-parameter language model into a compressor that emits both the medical jargon from the question and a question-aligned summary of five retrieved passages, and that this two-part output improves reader accuracy. During training, medical entities in a question are replaced with a special mask token, and the model is trained to generate, in order, the entity names, a short description of each, and the summary, with the retrieved passages providing the surrounding context that resolves the masks. At inference the same pipeline runs without any annotation, and the reader sees only the entity descriptions and the summary, not the retrieved documents. The paper reports consistent gains on MedQuAD, MASH-QA, and BioASQ across Llama-3-8B, Llama-3-70B, Mixtral-8x7B, GPT-4o, MedAlpaca-13B, and Meditron-70B readers, and it reports that the advantage persists on MEDIQA, a dataset the compressor never saw during training. Reranker preference and inference-speed data are offered as supporting evidence that the compressed context is more relevant to the question and cheaper to process than top-5 retrieval.
Load-bearing premise
The method assumes that the first sentence of each encyclopedia-style article is an accurate, useful description of the medical term in the question, the advantage over ordinary summarization largely disappears when those descriptions are withheld, and it evaluates only questions in which a named-entity recognizer finds at least one such term.
Editorial extensions
If this is right
- If the central claim holds, closed-domain retrieval-augmented QA can be improved without retraining or swapping the reader model: an entity-aware compressor that produces short glosses plus a concise summary beats raw retrieval and prior compressors across reader sizes.
- Readers that lack domain knowledge, the general-purpose LLMs in the study, benefit most from the injected descriptions, while medically trained readers benefit less, indicating the method works by supplying knowledge the reader is missing.
- Because the compressed prompt is roughly one-fifth the token length of five retrieved passages, the method about doubles reader throughput on the tested setup while improving accuracy, making it practical as a preprocessing step in RAG pipelines.
- The entity-focused training transfers to unseen data: performance on MEDIQA stays near the seen-data level even though the compressor was not trained on it, which the paper presents as evidence of usefulness in data-scarce closed domains.
Reading between the lines
- Extension: the entity-gloss mechanism likely transfers to other jargon-heavy closed domains, such as law, chemistry, or finance, wherever a retrieval corpus with title-and-first-sentence structure exists; the paper itself only claims the English biomedical setting.
- A testable consequence the paper leaves implicit: if gloss quality is the active ingredient, replacing the first-sentence glosses with richer or curated definitions should raise accuracy further, while deliberately wrong definitions should push scores toward the no-prior-knowledge ablation level.
- Scope boundary worth noting: the evaluation covers only questions in which the entity recognizer finds at least one medical term, with roughly a quarter to a third of each training dataset filtered out, so the reported gains apply to entity-bearing questions; a fallback for entity-free questions would be needed in a deployed system.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes K-COMP, a retrieval-augmented QA pipeline for the medical domain that compresses retrieved passages with a small fine-tuned language model (Gemma-2B). The compressor masks medical entities in the question, generates entity descriptions and a question-aligned summary under a causal masking objective, and appends the descriptions plus summary to the reader prompt. The authors evaluate on MedQuAD, MASH-QA, and BioASQ with six reader models, comparing against standard RAG, RECOMP, LLMLingua, and a summarization fine-tune (FineTune), and report wins on most table cells, plus a GPT-4o win-rate comparison on seen and unseen data.
Significance. If the causal knowledge-injection mechanism is genuinely responsible for the gains, K-COMP would be a practical and cost-effective method for closed-domain RAG: it uses a 2B compressor, shows solid gains over strong baselines (including RECOMP and FineTune) across many reader models, and appears to generalize to an unseen dataset (MEDIQA). The paper also ships code, evaluates six reader LLMs, and includes efficiency measurements, which are valuable reproducibility assets. However, the central mechanistic claim — that the causal masking objective injects useful domain knowledge — is not directly supported by the ablation evidence, and the lack of significance testing and the generic, unevaluated entity descriptions make the empirical contribution weaker than the headline claims suggest.
major comments (5)
- [§5.2, Table 2] The -Prior ablation does not support the central claim that the causal knowledge-injection objective produces the gains. Across many rows, K-COMP without the entity descriptions is statistically indistinguishable from, or slightly worse than, FineTune (e.g., MedQuAD Llama-3-8B BertScore 73.93 vs. 74.27, and MASH-QA Meditron-70B UniEval 60.42 vs. 60.06). Since -Prior and FineTune use the same compressor and differ only in the training objective, the observed advantage of K-COMP appears to come almost entirely from appending entity descriptions to the reader prompt — a prompt-formatting change that is independent of the causal masking training. The authors should either include FineTune with the same entity-description prompt, or directly evaluate summary quality (e.g., answer recall in the summary) to show that the objective matters beyond prompt formatting.
- [§4.2 and Table 14] The entity descriptions are assumed to be the first sentence of each MedCorp article, but their factual quality and task relevance are never evaluated. The qualitative examples show generic glosses such as "x-ray: Form of short-wavelength electromagnetic radiation" and "rheumatoid arthritis: Type of autoimmune arthritis," which are not specific prior knowledge needed to answer the given question. This undermines the claim that K-COMP "provides the knowledge required to answer correctly." The authors should assess description quality (e.g., through human annotation or automatic factuality checks) and demonstrate that the descriptions contain information needed for the QA task beyond what is already in the retrieved passages.
- [§5.2, Tables 1 and 2] No significance tests or error bars are reported anywhere. Many headline differences are 1–2 BertScore points (e.g., MASH-QA Mixtral UniEval 62.91 vs. RECOMP 62.35; MedQuAD Llama-3-70B BertScore 82.06 vs. RECOMP 79.75), and with a single run per condition and stochastic decoding, the reported ordering may not be stable. The authors should provide confidence intervals, paired significance tests across questions, or multiple seeds with variance reporting for at least the main comparison against FineTune and RECOMP.
- [§6.4 and Table 17] The unseen-data claim rests on only 140 MEDIQA questions, and the GPT-4o win-rate evaluation has no human agreement or alternative judge. Additionally, the compressor is trained on summaries generated by GPT-4o-mini and then judged by GPT-4o, which are from the same model family, so the evaluation may favor summaries with GPT-style phrasing. The authors should at least report the agreement of GPT-4o with human judges on a sample, or use a different reader/judge model for the win-rate comparison, and they should provide confidence intervals for the 140-question unseen evaluation.
- [§A.2, Table 15] The test sets are filtered to questions in which ScispaCy detects at least one entity and (for train/validation) for which a corpus description exists; this removes 5.2%–8.4% of test questions per dataset. This changes the test distribution relative to the unfiltered datasets on which baselines like Top-5 RAG are evaluated, so the reported numbers may not reflect performance on all medical questions. The authors should report results on the unfiltered test sets (or at least quantify how each baseline is affected by the filtering) to ensure a fair comparison.
minor comments (6)
- [Throughout] The paper has inconsistent naming: "K-comp", "K- COMP", "K-COMP", and "K- COMP" appear in the abstract and body; please standardize to "K-COMP".
- [§5.2, Table 2 caption] The caption says "−P rior" but the table header reads "− P rior"; please fix the spacing to a single consistent form.
- [§6.3, after Table 4] The sentence "FineTune merely summarises the passages...without considering the queried intent" contains a subject-verb agreement error; also "does not fully trusting" should be "does not fully trust".
- [§6.2, Table 3] The throughput claim of "double the throughput" is based on a single run on one GPU; please specify whether times are wall-clock or GPU-time and note the variance, since the claim is used to motivate efficiency.
- [§3] The phrase "regressively encoding the infilled span" is unclear; it should be "progressively" or "autoregressively" depending on intended meaning. The paper also uses "causal masking" without defining the exact attention mask; a short formal description would help.
- [References] The reference to "(Xu et al., 2023)" for the first-sentence-as-description assumption is a preprint (KILM); please update or clarify the provenance, since this assumption is load-bearing for the entity-description construction.
Circularity Check
No significant circularity: K-COMP is an empirical compression pipeline evaluated on independent benchmarks; the strongest validity concern is the −Prior ablation, which is a legitimate ablation rather than a circular reduction.
full rationale
This is an empirical systems paper, not a derivation chain. The central claim (Table 1) is an end-task comparison against RECOMP, LLMLingua, FineTune, and uncompressed RAG on three external benchmarks; these baselines and datasets are independent of the K-COMP training data and of the authors' prior work. The causal-masking objective (§3) is a fine-tuning loss, and its output is evaluated by reader models and rerankers in a 0-shot setting; there is no equation in the paper whose predicted quantity is defined from its own input. The strongest apparent concern—Table 2's −Prior ablation showing performance close to FineTune—is an ablation, not a circularity: it legitimately isolates the contribution of the entity descriptions, and the conclusion that the training objective contributes little beyond the descriptions is a validity threat, not a self-referential reduction. The entity descriptions are constructed from the same MedCorp retrieval corpus as the passages (first sentences of title/text pairs), so the 'prior knowledge' is not independent external knowledge; however, this is an assumption about knowledge independence, not a definitional equivalence between a claimed prediction and an input. The one self-citation (Sangwon Ryu et al., 2024, including author G. G. Lee) appears in the reference list and is not load-bearing. No fitted parameter is renamed as a prediction, and no uniqueness theorem from the authors' prior work is invoked. Under the stated rules, the paper shows no significant circularity; the appropriate score is 0.
Assumptions & free parameters
free parameters (2)
- Number of retrieved passages (|P|) =
5
- Entity description source =
First sentence of corpus article (Section 4.2)
assumptions (5)
- standard math Autoregressive next-token prediction with cross-entropy loss and CM3-style causal masking is a valid training objective for the compressor.
- domain assumption The first sentence of each MedCorp title-text pair is a correct short description of the entity and provides knowledge required to answer.
- domain assumption ScispaCy named-entity recognition reliably identifies all relevant medical entities in questions.
- domain assumption GPT-4o-mini generated summaries (which exclude the question) are high-quality supervision for the compressor.
- ad hoc to paper Excluding the question from the summary synthesis prompt produces better training targets for question-aligned compression.
Cite this review
Pith. "Pith review of K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor." pith.science (2026). https://pith.science/paper/O7MWWZ2K
@misc{pith2026250113567,
author = {Pith},
title = {Pith review of: K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor},
year = {2026},
howpublished = {\url{https://pith.science/paper/O7MWWZ2K}},
note = {Machine review of arXiv:2501.13567}
}
read the original abstract
Retrieval-augmented question answering (QA) integrates external information and thereby increases the QA accuracy of reader models that lack domain knowledge. However, documents retrieved for closed domains require high expertise, so the reader model may have difficulty fully comprehending the text. Moreover, the retrieved documents contain thousands of tokens, some unrelated to the question. As a result, the documents include some inaccurate information, which could lead the reader model to mistrust the passages and could result in hallucinations. To solve these problems, we propose K-comp (Knowledge-injected compressor) which provides the knowledge required to answer correctly. The compressor automatically generates the prior knowledge necessary to facilitate the answer process prior to compression of the retrieved passages. Subsequently, the passages are compressed autoregressively, with the generated knowledge being integrated into the compression process. This process ensures alignment between the question intent and the compressed context. By augmenting this prior knowledge and concise context, the reader models are guided toward relevant answers and trust the context.
Figures
Reference graph
Works this paper leans on
-
[1]
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, and Luke Zettlemoyer. 2022. https://arxiv.org/abs/2201.07520 Cm3: A causal masked multimodal model of the internet . Preprint, arXiv:2201.07520
arXiv 2022
-
[2]
Amin Ahmad, Noah Constant, Yinfei Yang, and Daniel Cer. 2019. https://doi.org/10.18653/v1/D19-5819 R e QA : An evaluation for end-to-end answer retrieval models . In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 137--146, Hong Kong, China. Association for Computational Linguistics
-
[3]
AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card
2024
-
[4]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024 a . https://openreview.net/forum?id=hSyW5go0v8 Self- RAG : Learning to retrieve, generate, and critique through self-reflection . In The Twelfth International Conference on Learning Representations
work page 2024
-
[5]
Akari Asai, Zexuan Zhong, Danqi Chen, Pang Wei Koh, Luke Zettlemoyer, Hannaneh Hajishirzi, and Wen tau Yih. 2024 b . https://arxiv.org/abs/2403.03187 Reliable, adaptable, and attributable language models with retrieval . Preprint, arXiv:2403.03187
arXiv 2024
-
[6]
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. 2022. https://arxiv.org/abs/2207.14255 Efficient training of language models to fill in the middle . Preprint, arXiv:2207.14255
arXiv 2022
-
[7]
Asma Ben Abacha and Dina Demner - Fushman. 2019. https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-019-3119-4 A question-entailment approach to question answering . BMC Bioinform. , 20(1):511:1--511:23
-
[8]
Asma Ben Abacha, Chaitanya Shivade, and Dina Demner-Fushman. 2019. https://doi.org/10.18653/v1/W19-5039 Overview of the MEDIQA 2019 shared task on textual inference, question entailment and question answering . In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 370--379, Florence, Italy. Association for Computational Linguistics
Show all 81 references
-
[9]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
2020
-
[10]
Zeming Chen, Alejandro Hern \'a ndez Cano, Angelika Romanou, Antoine Bonnet, Kyle Matoba, Francesco Salvi, Matteo Pagliardini, Simin Fan, Andreas K \"o pf, Amirkeivan Mohtashami, et al. 2023. Meditron-70b: Scaling medical pretraining for large language models. arXiv preprint a...
2023 arXiv
-
[11]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[12]
Chris Donahue, Mina Lee, and Percy Liang. 2020. https://doi.org/10.18653/v1/2020.acl-main.225 Enabling language models to fill in the blanks . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2492--2501, Online. Association for ...
2020 doi
-
[13]
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. https://doi.org/10.18653/v1/2022.acl-long.26 GLM : General language model pretraining with autoregressive blank infilling . In Proceedings of the 60th Annual Meeting of the Associatio...
2022 doi
-
[14]
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. https://openreview.net/forum?id=hQwb-lbM6EL Incoder: A generative model for code infilling and synthesis . In The Eleventh Internation...
2023
-
[15]
Giacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro, and Zaiqiao Meng. 2024. https://aclanthology.org/2024.acl-long.533 To generate or to retrieve? on the effectiveness of artificial contexts for medical open-domain question answering . In Proceedings of the 62nd A...
2024
-
[16]
Mandy Guo, Yinfei Yang, Daniel Cer, Qinlan Shen, and Noah Constant. 2021. https://aclanthology.org/2021.adaptnlp-1.10 M ulti R e QA : A cross-domain evaluation for R etrieval question answering models . In Proceedings of the Second Workshop on Domain Adaptation for NLP, pages ...
2021
-
[17]
Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander L \"o ser, Daniel Truhn, and Keno K Bressem. 2023. Medalpaca--an open-source collection of medical conversational ai models and training data. arXiv preprint arXiv:2304.08247
2023 arXiv
-
[18]
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. https://openreview.net/forum?id=rygGQyrFvH The curious case of neural text degeneration . In International Conference on Learning Representations
2020
-
[19]
Shengchao Hu, Li Shen, Ya Zhang, Yixin Chen, and Dacheng Tao. 2024. On transforming reinforcement learning with transformers: The development trajectory. IEEE Transactions on Pattern Analysis and Machine Intelligence
2024
-
[20]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276
2024 arXiv
-
[21]
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. https://arxiv.org/abs/2112.09118 Unsupervised dense information retrieval with contrastive learning . Preprint, arXiv:2112.09118
2022 arXiv
-
[22]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 a . https://doi.org/10.1145/3571730 Survey of hallucination in natural language generation . ACM Comput. Surv., 55(12)
2023 doi
-
[23]
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023 b . https://doi.org/10.18653/v1/2023.findings-emnlp.123 Towards mitigating LLM hallucination via self reflection . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 18...
2023 doi
-
[24]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088
2024 arXiv
-
[25]
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023 a . https://doi.org/10.18653/v1/2023.emnlp-main.825 LLML ingua: Compressing prompts for accelerated inference of large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natu...
2023 doi
-
[26]
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 b . https://doi.org/10.18653/v1/2023.emnlp-main.495 Active retrieval augmented generation . In Proceedings of the 2023 Conference on Empirical Methods...
2023 doi
-
[27]
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. https://doi.org/10.3390/app11146421 What disease does this patient have? a large-scale open domain question answering dataset from medical exams . Applied Sciences, 11(14)
2021 doi
-
[28]
Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Xiaojian Jiang, Jiexin Xu, Li Qiuxia, and Jun Zhao. 2024. https://aclanthology.org/2024.lrec-main.1466 Tug-of-war between knowledge: Exploring and resolving knowledge conflicts in retrieval-augmented language models . In Proceedin...
2024
-
[29]
Jeff Johnson, Matthijs Douze, and Herv \'e J \'e gou. 2019. Billion-scale similarity search with GPUs . IEEE Transactions on Big Data, 7(3):535--547
2019
-
[30]
Weld, Luke Zettlemoyer, and Omer Levy
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020. https://doi.org/10.1162/tacl_a_00300 S pan BERT : Improving pre-training by representing and predicting spans . Transactions of the Association for Computational Linguistics, 8:64--77
2020 doi
-
[31]
Minki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi, and Sung Ju Hwang. 2024. Knowledge-augmented reasoning distillation for small language models in knowledge-intensive tasks. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS...
2024
-
[32]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference on Empiric...
2020 doi
-
[33]
Jaehyung Kim, Jaehyun Nam, Sangwoo Mo, Jongjin Park, Sang-Woo Lee, Minjoon Seo, Jung-Woo Ha, and Jinwoo Shin. 2024. https://openreview.net/forum?id=w4DW6qkRmt Sure: Summarizing retrievals using answer candidates for open-domain QA of LLM s . In The Twelfth International Confer...
2024
-
[34]
Anastasia Krithara, Anastasios Nentidis, Konstantinos Bougiatiotis, and Georgios Paliouras. 2023. Bioasq-qa: A manually curated corpus for biomedical question answering. Scientific Data, 10(1):170
2023
-
[35]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. https://doi.org/10.1145/3600006.3613165 Efficient memory management for large language model serving with pagedattention . In Proceedings of the 29...
2023
-
[36]
Carlos Lassance and St\' e phane Clinchant. 2022. https://doi.org/10.1145/3477495.3531833 An efficiency study for splade models . In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22, page 2220–2226, New ...
2022
-
[37]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...
2020 doi
-
[38]
Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.391 Compressing context to enhance inference efficiency of large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, ...
2023 doi
-
[39]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[40]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024 a . https://proceedings.mlsys.org/paper_files/paper/2024/file/42a452cbafa9dd64e9ba4aa95cc1ef21-Paper-Conference.pdf Awq: Activation-aware w...
2024
-
[41]
Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Richard James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, and Wen tau Yih. 2024 b . https://openreview.net/forum?id=22OTbutug9 RA - DIT : Retrieval-augmented dual instruction ...
2024
-
[42]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. https://doi.org/10.1162/tacl_a_00638 Lost in the middle: How language models use long contexts . Transactions of the Association for Computational Linguistics, 12...
2024 doi
-
[43]
Shuai Liu, Hyundong Cho, Marjorie Freedman, Xuezhe Ma, and Jonathan May. 2023 a . https://doi.org/10.18653/v1/2023.acl-long.468 RECAP : Retrieval-enhanced context-aware prefix encoder for personalized dialogue response generation . In Proceedings of the 61st Annual Meeting of ...
2023 doi
-
[44]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 b . https://doi.org/10.18653/v1/2023.emnlp-main.153 G -eval: NLG evaluation using gpt-4 with better human alignment . In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...
2023 doi
-
[45]
Quanyu Long, Wenya Wang, and Sinno Pan. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.402 Adapt in contexts: Retrieval-augmented domain adaptation via in-context learning . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 652...
2023 doi
-
[46]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations
2019
-
[47]
Antoine Louis, Gijs van Dijck, and Gerasimos Spanakis. 2024. https://doi.org/10.1609/aaai.v38i20.30232 Interpretable long-form legal question answering with retrieval-augmented large language models . Proceedings of the AAAI Conference on Artificial Intelligence, 38(20):22266--22275
2024 doi
-
[48]
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. 2021. https://doi.org/10.18653/v1/2021.eacl-main.235 Entity-level factual consistency of abstractive text summarization . In Proceedings of the 16t...
2021 doi
-
[49]
Mark Neumann, Daniel King, Iz Beltagy, and Waleed Ammar. 2019. https://doi.org/10.18653/v1/W19-5034 S cispa C y: Fast and robust models for biomedical natural language processing . In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 319--327, Florence, Italy. Ass...
2019 doi
-
[50]
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. 2023. https://arxiv.org/abs/2303.13375 Capabilities of gpt-4 on medical challenge problems . Preprint, arXiv:2303.13375
2023 arXiv
-
[51]
Morris, Brandon Duderstadt, and Andriy Mulyar
Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. 2024. https://arxiv.org/abs/2402.01613 Nomic embed: Training a reproducible long context text embedder . Preprint, arXiv:2402.01613
2024 arXiv
-
[52]
OpenAI. 2024. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Gpt-4o mini: advancing cost-efficient intelligence
2024
-
[53]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. http://jmlr.org/papers/v21/20-074.html Exploring the limits of transfer learning with a unified text-to-text transformer . Journal of Machine Lea...
2020
-
[54]
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. https://doi.org/10.1162/tacl_a_00605 In-context retrieval-augmented language models . Transactions of the Association for Computational Linguistics, 11:1316--1331
2023 doi
-
[55]
Houxing Ren, Mingjie Zhan, Zhongyuan Wu, and Hongsheng Li. 2024. https://arxiv.org/abs/2405.17103 Empowering character-level text infilling by eliminating sub-tokens . Preprint, arXiv:2405.17103
2024 arXiv
-
[56]
Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval , 3(4):333--389
2009
-
[57]
Cheol Ryu, Seolhwa Lee, Subeen Pang, Chanyeol Choi, Hojun Choi, Myeonggee Min, and Jy-Yong Sohn. 2023. https://doi.org/10.18653/v1/2023.nllp-1.13 Retrieval-based evaluation for LLM s: A case study in K orean legal QA . In Proceedings of the Natural Legal Language Processing Wo...
2023 doi
-
[58]
Sangwon Ryu, Heejin Do, Yunsu Kim, Gary Geunbae Lee, and Jungseul Ok. 2024. https://arxiv.org/abs/2406.04625 Key-element-informed sllm tuning for document summarization . Preprint, arXiv:2406.04625
2024 arXiv
-
[59]
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024. https://openreview.net/forum?id=GN921JHCRw RAPTOR : Recursive abstractive processing for tree-organized retrieval . In The Twelfth International Conference on Learning Repres...
2024
-
[60]
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. https://doi.org/10.18653/v1/2024.naacl-long.463 REPLUG : Retrieval-augmented black-box language models . In Proceedings of the 2024 Conference of the Nor...
2024 doi
-
[61]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi \`e re, Mihir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295
2024 arXiv
-
[62]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[63]
Smith, Iz Beltagy, and Hannaneh Hajishirzi
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2024 a . How far can camels go? exploring the state of instruction tuning on open resources. In Pr...
2024
-
[64]
Yubo Wang, Xueguang Ma, and Wenhu Chen. 2024 b . https://arxiv.org/abs/2309.02233 Augmenting black-box llms with medical textbooks for clinical question answering . Preprint, arXiv:2309.02233
2024 arXiv
-
[65]
Zihao Wang, Anji Liu, Haowei Lin, Jiaqi Li, Xiaojian Ma, and Yitao Liang. 2024 c . https://arxiv.org/abs/2403.05313 Rat: Retrieval augmented thoughts elicit context-aware reasoning in long-horizon generation . Preprint, arXiv:2403.05313
2024 arXiv
-
[66]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[67]
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024. https://doi.org/10.1145/3626772.3657878 C-pack: Packed resources for general chinese embeddings . In Proceedings of the 47th International ACM SIGIR Conference on Research and Develop...
2024
-
[68]
Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024. Benchmarking retrieval-augmented generation for medicine. arXiv preprint arXiv:2402.13178
2024 arXiv
-
[69]
Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024 a . https://openreview.net/forum?id=mlJLVigNHp RECOMP : Improving retrieval-augmented LM s with context compression and selective augmentation . In The Twelfth International Conference on Learning Representations
2024
-
[70]
Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2024 b . https://arxiv.org/abs/2310.03025 Retrieval meets long context large language models . Preprint, arXiv:2310.03025
2024 arXiv
- [71]
-
[72]
Niraj Yagnik, Jay Jhaveri, Vivek Sharma, and Gabriel Pila. 2024. https://arxiv.org/abs/2401.11389 Medlm: Exploring language models for medical question answering systems . Preprint, arXiv:2401.11389
2024 arXiv
-
[73]
Haoyan Yang, Zhitao Li, Yong Zhang, Jianzong Wang, Ning Cheng, Ming Li, and Jing Xiao. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.326 PRCA : Fitting black-box large language models for retrieval question answering via pluggable reward-driven contextual adapter . In Proc...
2023 doi
-
[74]
Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023 a . https://openreview.net/forum?id=fB0hRu9GZUS Generate rather than retrieve: Large language models are strong context generators . In The Eleventh In...
2023
-
[75]
Wenhao Yu, Hongming Zhang, Xiaoman Pan, Kaixin Ma, Hongwei Wang, and Dong Yu. 2023 b . https://arxiv.org/abs/2311.09210 Chain-of-note: Enhancing robustness in retrieval-augmented language models . Preprint, arXiv:2311.09210
2023 arXiv
-
[76]
Zichun Yu, Chenyan Xiong, Shi Yu, and Zhiyuan Liu. 2023 c . https://doi.org/10.18653/v1/2023.acl-long.136 Augmentation-adapted retriever improves generalization of language models as generic plug-in . In Proceedings of the 61st Annual Meeting of the Association for Computation...
2023 doi
-
[77]
Weinberger, and Yoav Artzi
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with bert . In International Conference on Learning Representations
2020
-
[78]
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.131 Towards a unified multi-dimensional evaluator for text generation . In Proceedings of the 2022 Conference on Empiric...
2022 doi
-
[79]
Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K. Reddy. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.342 Question answering with long multiple-span answers . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3840--3849, Online...
2020 doi
-
[80]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[81]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.