REVIEW 3 major objections 5 minor 71 references
Investigating Task Arithmetic for Zero-Shot Information Retrieval
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper claims that adding the parameter difference between a domain-finetuned language model and its pretrained base to an IR-finetuned reranker produces a model that ranks better on unseen scientific, biomedical, and multilingual test…
desk verdict Solid multilingual zero-shot result, but the abstract's 'consistently improves' overstates the supporting evidence; the paper's own α=1 tables show mixed domain-transfer results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the task vector $\tau_D = \Theta_D - \Theta_0$, the parameter-wise difference between a domain-finetuned model and its pretrained base, which the paper treats as a portable encoding of domain shift. The adaptation identity is $\Theta' = \Theta_T + \alpha \tau_D$, where $\Theta_T$ is an MS-MARCO-finetuned reranker and $\alpha$ is a scalar controlling how much domain knowledge is injected. This machinery works because the paper assumes the domain shift and the ranking competence occupy compatible regions of parameter space, so a simple linear addition can inject the former without erasing the latter; the ablation shows $\alpha = 1.0$ is rarely optimal, so a small grid search over $\alpha$ is used to find the tradeoff between ranking competence and domain specialization.
What would settle it
Run the same six-model, eight-dataset protocol on a held-out domain not in the paper, such as legal or code retrieval, with a public $\Theta_D$ sharing the same $\Theta_0$: if no $\alpha \in [0.1, 1.0]$ yields NDCG@10 at or above the $\Theta_T$ baseline, the claimed consistent gains fail to generalize beyond the chosen domains.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that Task Arithmetic transfers domain and language competence into an IR model in parameter space: given a pretrained model $\Theta_0$, a domain-finetuned model $\Theta_D$, and an MS-MARCO-finetuned reranker $\Theta_T$, defining the task vector $\tau_D = \Theta_D - \Theta_0$ and forming $\Theta' = \Theta_T + \alpha \tau_D$ produces a reranker that, with an optimized $\alpha$, outperforms the strong $\Theta_T$ baselines across scientific (SciFact, SCIDOCS), biomedical (TREC-COVID, NFCorpus), and multilingual (GermanQuAD, MIRACL English, French, Spanish) test sets, with reported gains up to 18% in NDCG@10 and 15% in P@10. The paper also reports that in the fully zero-shot setting $\alpha = 1$ the gains are concentrated in the multilingual experiments; for the scientific and biomedical datasets the reliable gains require tuning $\alpha$ on development data.
Load-bearing premise
The core assumption is that a domain-specific shift learned under a language-modeling objective lives in the same parameter space as an MS-MARCO ranking model, so that simply adding the difference vector transfers domain knowledge without eroding ranking ability.
Editorial extensions
If this is right
- Any public domain- or language-finetuned model sharing a base with an IR reranker can be converted into a task vector and added to the reranker; no backpropagation or domain labels are needed.
- Multilingual reranking is the clearest beneficiary: with $\alpha = 1$, MT5-base adapted with language task vectors improves over the MS-MARCO-tuned baseline on German, Spanish, French, and English, including statistically significant gains up to 18% in NDCG@10.
- For scientific and biomedical retrieval, the gains are real but conditional: an $\alpha$ optimized on a small development set such as NFCorpus plus a 20% subset of SciFact training queries is needed, since $\alpha = 1$ often underperforms the IR baseline.
- Because the domain vector comes from a language-modeling or masked-language-modeling objective while the target model is trained for ranking, the transfer works across objectives, and optimal $\alpha$ values above 0.3 for all models indicate non-trivial domain knowledge is injected.
Reading between the lines
- A natural next test the paper does not run is adding multiple task vectors at once to see whether domain knowledge composes linearly when several specialties are merged into one reranker.
- The $\alpha$-sensitivity observed in the ablation suggests that a geometric predictor of optimal scaling, such as the norm or cosine similarity between $\tau_D$ and $\Theta_T$, could make the method truly zero-shot without a development set.
- In lower-resource or noisier domains, the quality of the public $\Theta_D$ model likely determines whether arithmetic transfer helps, so the reported gains should be read as upper bounds for well-behaved, openly available domain models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes applying task arithmetic to zero-shot information retrieval: given a pre-trained model Θ0, a domain-finetuned model ΘD, and an MS-MARCO-finetuned reranker ΘT, the method adds the domain task vector τD = ΘD − Θ0 to ΘT, scaled by α (Eq. 2), and evaluates the resulting reranker Θ′ on biomedical, scientific, and multilingual datasets. The authors report gains of up to 18% in NDCG@10 and 15% in P@10 and claim that Task Arithmetic consistently improves strong IR baselines without additional fine-tuning. The evaluation covers six model architectures and eight datasets, with statistical significance testing. The paper's own tables distinguish a fully zero-shot α=1 setting from an α-optimized setting using development data, and the multilingual experiments are conducted with α=1 only.
Significance. The underlying idea is timely and pragmatic: reusing publicly available domain-finetuned models through weight arithmetic could provide a training-free adaptation path for IR. The paper's strength is that the multilingual results with α=1 (Table 2) are statistically significant, genuinely zero-shot, and reproducible from public checkpoints with released code. However, the headline claim of consistent improvement is contradicted by the paper's own α=1 results on most biomedical/scientific configurations, and the domain gains rely on development-set selection of α and fusion weights. As reported, the evidence supports a narrower claim: task arithmetic helps zero-shot multilingual reranking, and helps biomedical/scientific reranking only when α is calibrated on labeled data. The paper also provides a useful ablation of α sensitivity, which is a positive analytical contribution.
major comments (3)
- [§4, Table 1; Abstract; Introduction] The abstract and Introduction state that Task Arithmetic "consistently improves upon strong IR baselines," but the paper's own α=1 results do not support this. Section 4 reports that with α=1 Task Arithmetic outperforms the MS-MARCO baselines on only four of the twenty model–dataset combinations in Table 1 (TREC-COVID with RoBERTa-base, T5-base, and T5-Large, and SCIDOCS with T5-Large), and it degrades performance on most others; for example, Llama-2 on SciFact drops from .770 to .757 NDCG@10, and DistilBERT on TREC-COVID drops from .744 to .675. The headline gains of up to 18% NDCG@10 come from the multilingual Table 2, where α=1 is genuinely zero-shot, or from the domain experiments where α is optimized on development sets. The central claim should be reframed to distinguish the zero-shot multilingual result from the development-calibrated domain result, or the "consistently improves" phrasing should be removed.
- [§3.2, Eq. (2); §4] The paper defines the fully zero-shot setting as α=1, but the experimental protocol for the biomedical/scientific domain includes an additional tuned component: the BM25/LLM fusion weights λ_BM25 and λ_LLM are optimized in [0,1] on the NFCorpus and SciFact development sets (Section 3.2). The reported domain gains for SciFact and NFCorpus therefore depend on two development-set hyperparameters, not just a single fixed α. To support the zero-shot characterization, the paper should report results with a fixed fusion rule (e.g., λ_BM25=λ_LLM=0.5) for all datasets, and clearly state which results use dev-optimized fusion weights.
- [§5, Table 3] The ablation directly undermines the practical zero-shot recommendation: Section 5 states that α=1.0 "rarely provides the best performance," and Table 3 shows large swings with α (e.g., T5-base on SciFact improves from .640 at α=1.0 to .722 at α=0.7, and DistilBERT drops from .723 at α=0.5 to .652 at α=1.0). Because no single α is consistently optimal across models or datasets, a practitioner operating without labels cannot choose it reliably. The paper should either provide a principled label-free selection rule for α, or explicitly restrict the zero-shot claim to settings where α=1 has been validated (i.e., the multilingual experiments).
minor comments (5)
- [Table 1 caption] The caption says "Best results are highlighted in boldface," but many bold entries carry no asterisk while some non-bold entries are marked with *; please clarify the relationship between bold highlighting and the significance markers.
- [Abstract] The abstract reports "gains of up to 18% in NDCG@10 and 15% in P@10" without specifying that these are relative gains over the MS-MARCO-tuned baseline in the multilingual setting; please add a qualifier such as "relative to the IR-tuned baseline in multilingual zero-shot evaluation."
- [Introduction / Conclusion] The phrase "consistently improves" appears in both the Introduction and the Conclusion; if the framing is revised per the major comments, both occurrences should be updated consistently.
- [Reference [7]] Reference [7] contains an extra comma in the author list ("Marzieh Fadaee, , Roberto Lotufo"); please fix the formatting.
- [§3.1] The paper excludes Wikipedia-based BEIR datasets on the grounds that the pretrained models have seen Wikipedia; this rationale should be discussed as a limitation, since it restricts the evaluation to domains where compatible domain-finetuned models exist and limits the generality of the "across the board" conclusions.
Circularity Check
No circularity: task-vector construction is an empirical hypothesis evaluated on external benchmarks, and optimized-α results are explicitly labeled as such.
full rationale
The paper's derivation chain is Θ′ = Θ_T + α(Θ_D − Θ_0), with τ_D = Θ_D − Θ_0 (Eqs. 1–2). This is not a self-definitional or input-equivalent construction: the task vectors are computed from publicly released, independently trained checkpoints, and the evaluation measures retrieval effectiveness on external BEIR and MIRACL datasets. No fitted parameter is renamed as a prediction: Section 2.2 explicitly distinguishes α = 1 (zero-shot) from α optimized on development data, and Table 1 labels the optimized rows as such. The paper's self-citations (e.g., refs. 5, 10–13, 31) are incidental related-work references and do not carry the load-bearing premise; there is no imported uniqueness theorem or ansatz hidden in a citation. The manuscript itself flags the key limitation: Section 4 states that with α = 1, Task Arithmetic outperforms the MS-MARCO baselines on only four of twenty model–dataset combinations, and Section 5 states that 'α = 1.0 rarely provides the best performance.' These passages weaken the headline 'zero-shot consistently improves' claim as a rigor matter, but they are honest negative results, not signs that the positive results are forced by construction. The optimized-α numbers for NFCorpus and SciFact are selected using those datasets' development splits, so they should be read as calibrated rather than fully zero-shot evidence, but this is an evaluation-transparency concern, not circular reasoning. The central claim has independent empirical content and is not equivalent to its inputs by definition.
Assumptions & free parameters
free parameters (2)
- α task-vector scaling factor =
0.3 to 1.0, depending on model and dataset (e.g., 0.8 for Llama-2, 0.7 for T5-base, 0.3 for RoBERTa-base)
- λ_BM25 and λ_LLM fusion weights =
Optimized in [0,1] on NFCorpus and SciFact development sets; 0.5 and 0.5 for other datasets
assumptions (4)
- domain assumption A task vector computed from language-model or masked-language-model fine-tuning can be added to an IR-tuned model with the same pretrained initialization without destroying retrieval competence.
- domain assumption The domain-fine-tuned model and the IR-fine-tuned model share the exact same architecture and pretrained weights Θ0.
- domain assumption BM25 top-100 candidate generation followed by re-ranking with a fixed fusion rule is a fair evaluation protocol.
- ad hoc to paper Wikipedia-based BEIR datasets are excluded because pretrained models already saw Wikipedia during pretraining.
Cite this review
Pith. "Pith review of Investigating Task Arithmetic for Zero-Shot Information Retrieval." pith.science (2026). https://pith.science/paper/SAFEN45R
@misc{pith2026250500649,
author = {Pith},
title = {Pith review of: Investigating Task Arithmetic for Zero-Shot Information Retrieval},
year = {2026},
howpublished = {\url{https://pith.science/paper/SAFEN45R}},
note = {Machine review of arXiv:2505.00649}
}
read the original abstract
Large Language Models (LLMs) have shown impressive zero-shot performance across a variety of Natural Language Processing tasks, including document re-ranking. However, their effectiveness degrades on unseen tasks and domains, largely due to shifts in vocabulary and word distributions. In this paper, we investigate Task Arithmetic, a technique that combines the weights of LLMs pre-trained on different tasks or domains via simple mathematical operations, such as addition or subtraction, to adapt retrieval models without requiring additional fine-tuning. Our method is able to synthesize diverse tasks and domain knowledge into a single model, enabling effective zero-shot adaptation in different retrieval contexts. Extensive experiments on publicly available scientific, biomedical, and multilingual datasets show that our method improves state-of-the-art re-ranking performance by up to 18% in NDCG@10 and 15% in P@10. In addition to these empirical gains, our analysis provides insights into the strengths and limitations of Task Arithmetic as a practical strategy for zero-shot learning and model adaptation. We make our code publicly available at https://github.com/DetectiveMB/Task-Arithmetic-for-ZS-IR.
Figures
Reference graph
Works this paper leans on
-
[1]
Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, and David Sontag
-
[2]
Samuel Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2023. Git Re- Basin: Merging Models modulo Permutation Symmetries. In The Eleventh Inter- national Conference on Learning Representations
work page 2023
-
[3]
Tiago Almeida and Sérgio Matos. 2024. Exploring efficient zero-shot synthetic dataset generation for Information Retrieval. In Findings of the Association for Computational Linguistics: EACL 2024, Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian’s, Malta, 1214–1231. https: //aclanthology.org/2024.findings-eacl.81/
work page 2024
-
[4]
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the Cross-lingual Transferability of Monolingual Representations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 4623–4637. https:...
-
[5]
Elias Bassani, Pranav Kasela, Alessandro Raganato, and Gabriella Pasi. 2022. A multi-domain benchmark for personalized search evaluation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 3822–3827
work page 2022
-
[6]
Rishabh Bhardwaj, Duc Anh Do, and Soujanya Poria. 2024. Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Comp...
-
[7]
Luiz Henrique Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, , Roberto Lotufo, and Rodrigo Nogueira. 2021. mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset. arXiv:2108.13897 [cs.CL]
arXiv 2021
-
[9]
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A full-text learning to rank dataset for medical information retrieval. InAdvances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20–23, 2016. Proceedings 38 . Springer, 716–722
2016
Show all 71 references
-
[10]
Marco Braga. 2024. Personalized Large Language Models through Parameter Efficient Fine-Tuning Techniques. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Wash- ington DC, USA) (SIGIR ’24). Association for Compu...
2024
-
[11]
Marco Braga, Pranav Kasela, Alessandro Raganato, and Gabriella Pasi. 2024. Synthetic Data Generation with Large Language Models for Personalized Com- munity Question Answering. In 2024 IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT)
2024
-
[12]
Marco Braga, Alessandro Raganato, and Gabriella Pasi. 2024. AdaKron: An Adapter-based Parameter Efficient Model Tuning with Kronecker Product. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING...
2024
-
[13]
Marco Braga, Alessandro Raganato, Gabriella Pasi, et al. 2023. Personalization in BERT with Adapter Modules and Topic Modelling. In Proceedings of the 13th Italian Information Retrieval Workshop (IIR 2023). Pisa, Italy . 24–29
2023
-
[14]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
2020
-
[15]
Rémi Calizzano, Malte Ostendorff, Qian Ruan, and Georg Rehm. 2022. Gen- erating Extended and Multilingual Summaries with Pre-trained Transformers. In Proceedings of the Thirteenth Language Resources and Evaluation Conference , Nicoletta Calzolari, Frédéric Béchet, Philippe Bla...
2022
-
[16]
Xinran Chen, Xuanang Chen, Ben He, Tengfei Wen, and Le Sun. 2024. Analyze, Generate and Refine: Query Expansion with LLMs for Zero-Shot Open-Domain QA. In Findings of the Association for Computational Linguistics: ACL 2024 , Lun- Wei Ku, Andre Martins, and Vivek Srikumar (Eds....
2024 doi
-
[17]
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz. 2022. Fusing finetuned models for better pretraining. arXiv preprint arXiv:2204.03044 (2022)
2022 arXiv
-
[18]
Alexandra Chronopoulou, Jonas Pfeiffer, Joshua Maynez, Xinyi Wang, Sebas- tian Ruder, and Priyanka Agrawal. 2024. Language and Task Arithmetic with Parameter-Efficient Layers for Zero-Shot Summarization. In Proceedings of the Fourth Workshop on Multilingual Representation Lear...
2024 doi
-
[19]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld
-
[20]
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, and Edoardo Ponti. 2024. Elastic Weight Removal for Faithful and Abstractive Dialogue Gen- eration. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: ...
2024 doi
-
[21]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human...
2019
-
[22]
Schwab, and Ari S
Jonathan Frankle, David J. Schwab, and Ari S. Morcos. 2020. The Early Phase of Neural Network Training. InInternational Conference on Learning Representations. https://openreview.net/forum?id=Hkl1iRNFwS
2020
-
[23]
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. In Proceedings of ACL
2020
-
[24]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International conference on machine learning. PMLR, 2790–2799
2019
-
[25]
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
2021
-
[26]
Shih-Cheng Huang, Pin-Zu Li, Yu-chi Hsu, Kuang-Ming Chen, Yu Tung Lin, Shih-Kai Hsiao, Richard Tsai, and Hung-yi Lee. 2024. Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages. In Proceedings of the 62nd Annual Meeting o...
2024
-
[27]
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Han- naneh Hajishirzi, and Ali Farhadi. 2022. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations
2022
-
[28]
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and An- drew Gordon Wilson. 2018. Averaging weights leads to wider optima and better generalization. In 34th Conference on Uncertainty in Artificial Intelligence 2018, UAI 2018. Association For Uncertainty in A...
2018
-
[29]
Shashank Mohan Jain. 2022. Hugging face. In Introduction to transformers for NLP: With the hugging face library and models to solve problems . Springer, 51–67
2022
-
[31]
Pranav Kasela, Marco Braga, Gabriella Pasi, and Raffaele Perego. 2024. SE-PQA: Personalized Community Question Answering. In Companion Proceedings of the ACM Web Conference 2024 (Singapore, Singapore) (WWW ’24). Association for Computing Machinery, New York, NY, USA, 1095–1098...
2024
-
[32]
Pranav Kasela, Gabriella Pasi, Raffaele Perego, and Nicola Tonellotto. 2024. Desire- me: Domain-enhanced supervised information retrieval using mixture-of-experts. In European Conference on Information Retrieval . Springer, 111–125. SIGIR ’25, July 13–18, 2025, Padua, Italy Ma...
2024
-
[33]
Smith, and Luke Zettlemoyer
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A. Smith, and Luke Zettlemoyer. 2022. Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models. In First Workshop on Interpolation Regular- izers and Beyond at NeurIPS 2022 . http...
2022
-
[34]
Minghan Li, Honglei Zhuang, Kai Hui, Zhen Qin, Jimmy Lin, Rolf Jagerman, Xuanhui Wang, and Michael Bendersky. 2024. Can Query Expansion Improve Generalization of Strong Cross-Encoder Rankers?. In Proceedings of the 47th International ACM SIGIR Conference on Research and Develo...
2024
-
[35]
Jimmy Lin, Rodrigo Nogueira, and Andrew Yates. 2022. Pretrained transformers for text ranking: Bert and beyond . Springer Nature
2022
-
[36]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL] https://arxiv.org/abs/1907.11692
2019 arXiv
-
[37]
Yuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Scott Yih, and Madian Khabsa. 2022. UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics...
2022
-
[38]
Michael S Matena and Colin A Raffel. 2022. Merging models with fisher-weighted averaging. Advances in Neural Information Processing Systems 35 (2022), 17703– 17716
2022
-
[39]
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023. Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey. ACM Comput. Surv. 56, 2, Article 30 (Sept...
2023 doi
-
[40]
Timo Möller, Julian Risch, and Malte Pietsch. 2021. GermanQuAD and Ger- manDPR: Improving Non-English Question Answering and Passage Retrieval. In Proceedings of the 3rd Workshop on Machine Reading for Question Answering . 42–50
2021
-
[41]
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cogni- tive Computation: Integrating neural and symbolic approaches 2016 ...
2016
-
[42]
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Docu- ment Ranking with a Pretrained Sequence-to-Sequence Model. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computat...
2020 doi
-
[43]
Marinela Parović, Ivan Vulić, and Anna Korhonen. 2024. Investigating the Potential of Task Arithmetic for Cross-Lingual Transfer. In Proceedings of the 18th Conference of the European Chapter of the Association for Computa- tional Linguistics (Volume 2: Short Papers) , Yvette ...
2024
-
[44]
Long N Phan, James T Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, and Grégoire Altan-Bonnet. 2021. Scifive: a text-to-text trans- former model for biomedical literature. arXiv preprint arXiv:2106.03598 (2021)
2021 arXiv
-
[45]
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Ben- dersky. 2024. Large Language Models are Effective Text Rankers with Pair- wise Ranking Prompting. In Findings of the Associat...
2024 doi
-
[46]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67
2020
-
[47]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing . Association for Computational Linguistics. http://arxiv.org/abs/1908.10084
2019 arXiv
-
[48]
Daniele Rizzo, Alessandro Raganato, and Marco Viviani. 2024. Comparatively Assessing Large Language Models for Query Expansion in Information Retrieval via Zero-Shot and Chain-of-Thought Prompting. InProceedings of the 14th Italian Information Retrieval Workshop, Udine, Italy,...
2024
-
[49]
Anna Rogers and Sasha Luccioni. 2024. Position: Key Claims in LLM Research Have a Long Tail of Footnotes. In Forty-first International Conference on Machine Learning
2024
-
[50]
Omid Rohanian, Mohammadmahdi Nouriborji, Samaneh Kouchaki, and David A Clifton. 2023. On the effectiveness of compact biomedical transformers. Bioinfor- matics 39, 3 (2023), btad103
2023
-
[51]
Omid Rohanian, Mohammadmahdi Nouriborji, Samaneh Kouchaki, Farhad Nooralahzadeh, Lei Clifton, and David A Clifton. 2024. Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing. Artificial Intelligence in Medicine 158 (2024), 103007. https://doi.org...
2024
-
[52]
V Sanh. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.. In Proceedings of Thirty-third Conference on Neural Information Processing Systems (NIPS2019)
2019
-
[53]
Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma, Hu Xu, Xi Victoria Lin, Baptiste Rozière, Jacob Kahn, Daniel Li, Wen-tau Yih, Jason Weston, et al. 2024. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM. arXiv preprint arXiv:2403.07816 (2024)
2024 arXiv
-
[54]
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Tra...
2021
-
[55]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, and Yasmine Babaei et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288 [cs.CL] https://arxiv.org/abs/2307.09288
2023 arXiv
-
[56]
Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R. Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. TREC-COVID: constructing a pandemic information retrieval test collection. SI- GIR Forum 54, 1, Article 1 (Feb. 2021), 12 pages. ht...
2021 doi
-
[57]
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7534–7550
2020
-
[58]
Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin, Wang Lu, Yiqiang Chen, Wenjun Zeng, and Philip S. Yu. 2023. Generalizing to Unseen Domains: A Survey on Domain Generalization. IEEE Transactions on Knowledge and Data Engineering 35, 8 (2023), 8052–8072. https://doi...
2023
-
[59]
Shuai Wang, Harrisen Scells, Shengyao Zhuang, Martin Potthast, Bevan Koopman, and Guido Zuccon. 2024. Zero-shot Generative Large Language Models for Systematic Review Screening Automation. InEuropean Conference on Information Retrieval. Springer, 403–420
2024
-
[60]
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Si- mon Kornblith, et al. 2022. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increas...
2022
-
[61]
Yuchen Xia, Jiho Kim, Yuhan Chen, Haojie Ye, Souvik Kundu, Cong Callie Hao, and Nishil Talati. 2024. Understanding the performance and estimating the cost of llm fine-tuning. In 2024 IEEE International Symposium on Workload Characteri- zation (IISWC). IEEE, 210–223
2024
-
[62]
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. arXiv:2010.11934 [cs.CL] https://arxiv.org/ abs/2010.11934
2021 arXiv
-
[63]
Prateek Yadav, Derek Tam, Leshem Choshen, Colin A Raffel, and Mohit Bansal
-
[64]
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024. Language models are super mario: Absorbing abilities from homologous models as a free lunch. In Forty-first International Conference on Machine Learning
2024
-
[65]
Jingtao Zhan, Qingyao Ai, Yiqun Liu, Jiaxin Mao, Xiaohui Xie, Min Zhang, and Shaoping Ma. 2022. Disentangled modeling of domain and relevance for adaptable dense retrieval. arXiv preprint arXiv:2208.05753 (2022)
2022 arXiv
-
[66]
Longhui Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, and Min Zhang. 2023. RankingGPT: Empowering Large Language Models in Text Ranking with Progressive Enhancement. arXiv:2311.16720 [cs.IR]
2023 arXiv
-
[67]
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Al- fonso Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin
-
[68]
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Haonan Chen, Zheng Liu, Zhicheng Dou, and Ji-Rong Wen. 2023. Large lan- guage models for information retrieval: A survey. arXiv preprint arXiv:2308.07107 (2023)
2023
-
[69]
Shengyao Zhuang, Honglei Zhuang, Bevan Koopman, and Guido Zuccon. 2024. A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information ...
2024
-
[71]
Transactions of the Association for Computational Linguistics 11 (2023), 1114–1131
MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages. Transactions of the Association for Computational Linguistics 11 (2023), 1114–1131
2023
-
[2020]
In Proceedings of the 58th Annual Meeting of the As- sociation for Computational Linguistics
SPECTER: Document-level Representation Learning using Citation- informed Transformers. In Proceedings of the 58th Annual Meeting of the As- sociation for Computational Linguistics . 2270–2282
-
[2022]
In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.)
Large language models are few-shot clinical information extractors. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United A...
2022 doi
-
[2023]
Advances in Neural Information Processing Systems 36 (2023), 7093–7115
Ties-merging: Resolving interference when merging models. Advances in Neural Information Processing Systems 36 (2023), 7093–7115
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.