REVIEW 5 major objections 5 minor 75 references
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read LUSIFER shows that an English-only-trained connector can make an English-centric LLM embedding model competitive across 14 languages, beating E5-Mistral by 3.19 points.
desk verdict Solid English-only recipe for multilingual adaptation of LLM embedders, but the SOTA claim is oversold and the benchmark needs transparency. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is the connector: a two-layer feed-forward network that maps XLM-R's hidden states to Mistral's hidden dimension, plus one trainable token appended to the sequence, a pattern borrowed from multimodal alignment. Training has two stages, both on English data: stage one aligns encoder and connector with the frozen LLM using masked reconstruction and autoregressive completion; stage two finetunes everything end-to-end with contrastive learning, in-batch and hard negatives, bidirectional attention, and LoRA. The connector's job is to make the English-centric LLM interpret XLM-R's language-neutral vectors as if they were its own native text, so multilingual semantics transfer without multilingual supervision.
What would settle it
Train the identical connector and contrastive pipeline but replace XLM-R with a monolingual English encoder, then evaluate on the same 14-language benchmark; if the multilingual advantage over the English-only baseline disappears or reverses, the claimed transfer from a language-universal space is not what drives the gains.
Extended reading notes
Core claim
The paper's central claim is that the language-agnostic representations of XLM-R can be projected into the input space of an English-centric LLM (Mistral-7B) through a two-layer feed-forward connector plus a single trainable token, and that after two stages of English-only training the LLM produces strong embeddings for languages it rarely encountered during pretraining. The authors report state-of-the-art results on 10 of 14 languages in their benchmark, an average of 62.63 versus E5-Mistral's 59.44, and a cross-lingual average of 57.89 versus 52.14, including a near-doubling on the IndicCrosslingual dataset. They also show through t-SNE that LUSIFER's representations mix languages far more than E5-Mistral's, which they take as evidence of the intended language-agnostic behavior.
Load-bearing premise
The approach assumes that XLM-R's hidden representations are already language-agnostic enough that a connector trained only on English text can teach an English-centric LLM to correctly interpret them for any language.
Editorial extensions
If this is right
- English-only training is enough to produce competitive multilingual embeddings across classification, clustering, reranking, retrieval, and STS, reducing the need for expensive multilingual corpora.
- A small connector plus public English datasets can come within two points of large multilingual embedding models that required extensive multilingual training data.
- Low-resource languages covered by XLM-R can see large embedding-quality gains, up to 22.15 points for Telugu over E5-Mistral.
- Cross-lingual retrieval, where queries and documents are in different languages, improves by 5.75 points on average over the strongest English-centric baseline, with the largest gains on Indic languages.
- The new 123-dataset, 14-language benchmark provides a common evaluation surface for future multilingual embedding research.
Reading between the lines
- Beyond the paper: the connector recipe should transfer to other encoder-LLM pairs, so replacing XLM-R or Mistral with larger or instruction-tuned counterparts could scale the gains further.
- Beyond the paper: because alignment is purely generative, the method inherits the multilingual coverage of the encoder; languages absent from XLM-R's pretraining would need encoder extensions before the connector can help.
- Beyond the paper: task-level results in the paper show reranking lags the strongest baseline, so practitioners should check task-specific trade-offs before deploying the model in production multilingual search.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LUSIFER, a two-stage method for adapting English-centric LLM embedding models to multilingual text embedding tasks without multilingual supervision for the embedding task itself. The architecture combines XLM-R-large as a multilingual encoder, a two-layer feed-forward connector with one trainable token, and Mistral-7B as the target LLM; the connector is first aligned with English masked-reconstruction and autoregressive objectives, and then all components are fine-tuned with contrastive learning on English retrieval data using LoRA. The authors evaluate on a newly assembled benchmark of 123 datasets across 14 languages and five embedding tasks, report improvements over the English-centric E5-Mistral model in 10 of 14 languages, report cross-lingual results on five datasets, and include ablations and a t-SNE visualization.
Significance. If the central claim were fully supported, the paper would demonstrate a practical recipe: a small connector trained only on English can transfer multilingual understanding from XLM-R to a large English-centric LLM, yielding large gains in medium- and low-resource languages. The paper has several strengths: the architecture is simple and clearly described, the two-stage training pipeline and hyperparameters are specified in detail, English-only training data is enumerated, ablations isolate the contributions of the alignment and fine-tuning stages, and the authors provide a code and data link. However, the evaluation currently rests on a self-constructed benchmark whose composition is not fully disclosed, and the headline 'state-of-the-art in 10 out of 14 languages' claim is only valid relative to E5-Mistral, not against the multilingual baselines included in the same table. Because these issues affect the central empirical claim, they need to be resolved before the paper can be recommended for acceptance.
major comments (5)
- [Section 4.3, Table 1] The statement that LUSIFER 'achieves state-of-the-art performance in 10 out of 14 languages' is only true when the comparison is restricted to E5-Mistral. Against the full baseline set in Table 1, LUSIFER is the highest-scoring model in only one language (Indonesian); BGE-M3 outperforms LUSIFER in 11 of 14 languages and Multilingual-E5-large outperforms it in 11 of 14. The average score of 62.63 is also lower than BGE-M3 (64.80) and Multilingual-E5-large (64.54). The state-of-the-art claim as written is therefore unsupported and should be revised to refer specifically to English-centric LLM baselines, not to the general class of multilingual embedding models.
- [Section 4.1, Tables 1 and 7-16] The aggregate comparison is difficult to interpret because the benchmark averages macro over languages with highly unequal numbers of datasets. The paper does not provide a per-language dataset count table, and the appendix shows that some languages have very few datasets (e.g., Farsi has six, and Telugu and Swahili have very few). The reported 3.19-point gain over E5-Mistral is therefore heavily driven by low-resource languages where E5-Mistral is expected to be weak, which does not isolate the connector's transfer mechanism. The authors should release the full dataset-to-language mapping and report per-language and per-dataset counts, ideally with standard errors or a sensitivity analysis over benchmark composition.
- [Section 4.1, Table 2, Abstract] The abstract states that cross-lingual evaluation uses four datasets, while Section 4.1 and Table 2 list five: Belebele, MLQA, STS17, STS22, and IndicCrosslingual. Additionally, the exact assignment of datasets to the 14 benchmark languages is not provided in the paper or in the linked repository description. This makes the central benchmark claim non-reproducible and must be fixed before the empirical results can be assessed.
- [Abstract, Section 1, Section 4.2] The claim that LUSIFER works 'without requiring explicit multilingual training data' is overstated. The method relies on XLM-R-large, which is pretrained on multilingual corpora; what the training pipeline avoids is explicit multilingual supervision for the embedding task. This distinction matters because the transfer mechanism may derive from XLM-R's multilingual pretraining rather than from the connector itself. The wording should be corrected throughout, including the abstract, to 'without multilingual supervision for the embedding task.'
- [Section 4.7, Figure 6] The evidence for the central 'language-agnostic universal space' hypothesis is a qualitative t-SNE plot of 200 samples from SIB200. This does not directly test whether the connector transfers multilingual semantics to the target LLM, and it is not supported by a quantitative metric. The Frozen Multilingual Encoder ablation in Table 5 shows that fine-tuning XLM-R matters, but a more direct test, such as comparing LUSIFER against running the same contrastive fine-tuning directly on XLM-R embeddings or measuring language-identification accuracy of the output embeddings, would strengthen the mechanism claim.
minor comments (5)
- [Table 5] The text says the Frozen Multilingual Encoder variant scored 56.74, but Table 5 reports 58.74; also, the sentence 'showed a 18.45 point' is missing the word 'drop' and should be reworded.
- [Figures 3-5] The per-dataset labels in the task-specific comparison figures are very small and difficult to read; the figures should be reformatted or supplemented with a table of per-task averages.
- [Section 3.1, Eq. (1)] The aligned hidden state equation is not numbered and the masking mechanism for padding tokens is described only in prose; adding a formal definition would improve clarity.
- [Section 4.1] The statement that the benchmark contains 123 datasets should clarify whether this counts dataset-language instances or unique datasets; the sum 48+24+24+22+5=123 suggests the former, and this should be stated explicitly.
- [Table 3] The hyperparameter table does not specify which modules LoRA is applied to (for example, attention projections only or also feed-forward layers); this detail is needed for reproducibility.
Circularity Check
No significant circularity: LUSIFER's multilingual gains are evaluated on external benchmarks and do not reduce to the paper's own definitions or fitted targets.
full rationale
LUSIFER's central claim is an empirical transfer result: a two-layer connector plus LoRA finetuning on English-only data yields multilingual embedding gains. The evaluation is entirely against pre-existing datasets and external baselines. Nothing in Section 3 defines the target metric in terms of the model's own parameters: Stage 1 loss (masked LM + next-token prediction) and Stage 2 contrastive loss are trained on English corpora and evaluated on held-out multilingual and cross-lingual datasets. No multilingual evaluation score enters the training objective, and no parameter is fitted to the multilingual test sets. The 'language-universal space' is presented as a hypothesis supported by external citations [30,46] and by t-SNE, not as a theorem that the results are derived from. The only self-citations are [24] for the high/medium/low-resource language grouping and [37] in a list of prior embedding work; neither supplies a conclusion that the paper's gains reduce to. The paper does contain a claim-calibration issue worth noting as correctness risk, not circularity: Section 4.3 says 'state-of-the-art performance in 10 out of 14 languages' while Table 1 shows BGE-M3 and Multilingual-E5-large outperform LUSIFER in 11 and 12 of the 14 languages respectively; this affects the strength of the SOTA claim, but it is an internal-consistency/benchmark-composition concern rather than a circular reduction. Accordingly, no circular step meets the evidence bar.
Assumptions & free parameters
free parameters (5)
- mask_ratio =
0.5
- num_hard_negatives =
7
- lora_rank =
16
- lora_alpha =
32
- connector_layers =
2
assumptions (4)
- domain assumption XLM-R representations are language-agnostic enough for zero-shot transfer to other languages
- ad hoc to paper A connector trained only on English can align multilingual encoder outputs with the target LLM's space
- domain assumption The 123-dataset, 14-language benchmark is a valid and unbiased measure of multilingual embedding quality
- domain assumption Macro-averaging language scores gives a meaningful overall performance metric
invented entities (2)
-
trainable token t
-
language-agnostic universal space
Cite this review
Pith. "Pith review of LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models." pith.science (2026). https://pith.science/paper/HA2AZADK
@misc{pith2026250100874,
author = {Pith},
title = {Pith review of: LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/HA2AZADK}},
note = {Machine review of arXiv:2501.00874}
}
read the original abstract
Recent advancements in large language models (LLMs) based embedding models have established new state-of-the-art benchmarks for text embedding tasks, particularly in dense vector-based retrieval. However, these models predominantly focus on English, leaving multilingual embedding capabilities largely unexplored. To address this limitation, we present LUSIFER, a novel zero-shot approach that adapts LLM-based embedding models for multilingual tasks without requiring multilingual supervision. LUSIFER's architecture combines a multilingual encoder, serving as a language-universal learner, with an LLM-based embedding model optimized for embedding-specific tasks. These components are seamlessly integrated through a minimal set of trainable parameters that act as a connector, effectively transferring the multilingual encoder's language understanding capabilities to the specialized embedding model. Additionally, to comprehensively evaluate multilingual embedding performance, we introduce a new benchmark encompassing 5 primary embedding tasks, 123 diverse datasets, and coverage across 14 languages. Extensive experimental results demonstrate that LUSIFER significantly enhances the multilingual performance across various embedding tasks, particularly for medium and low-resource languages, without requiring explicit multilingual training data.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Alabi, Yanke Mao, Haonan Gao, and Annie En-Shiun Lee
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, and Annie En-Shiun Lee. 2024. SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects. arXiv:2309.07445 [cs.CL] https://arxiv.org/abs/2309.07445
arXiv 2024
-
[3]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikoł aj Bińko...
-
[4]
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the Cross-lingual Transferability of Monolingual Representations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 4623–4637. doi:10...
-
[5]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https://arxiv.org/abs/1611.09268
arXiv 2018
-
[6]
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele Benchmark: a Parallel Reading Compre- hension Dataset in 122 Language Variants. InProceedings of the 62nd Annual Meet- ing of the Association for Computational Linguistic...
-
[7]
Rachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta, Sriram Ganapa- thy, Abhishek Bapna, Prateek Jain, and Partha Talukdar. 2024. LLM Augmented LLMs: Expanding Capabilities through Composition. In The Twelfth Interna- tional Conference on Learning Representations . https://openreview.net/forum?id= jjA4O1vJRz
work page 2024
-
[8]
Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. arXiv:2404.05961 [cs.CL] https://arxiv.org/abs/ 2404.05961
arXiv 2024
-
[9]
Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Man- ning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Pro- cessing, Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association for Computational Linguistics, Lisbon, Por...
Show all 75 references
-
[10]
Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu
-
[11]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil- laume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learn- ing at Scale. In Proceedings of the 58th Annual Me...
2020 doi
-
[12]
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bow- man, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating Cross- lingual Sentence Representations. In Proceedings of the 2018 Conference on Em- pirical Methods in Natural Language Processing , E...
2018 doi
-
[13]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2019
-
[14]
2024.PyTorch Lightning
William Falcon and The PyTorch Lightning team. 2024.PyTorch Lightning. doi:10. 5281/zenodo.13254264
2024
-
[15]
Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan. 2021. Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021) , Anna SIGIR ’25, July 13–18, 2025, Padua, Italy Hieu Man, ...
2021
-
[16]
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih ...
2021 doi
-
[17]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] https://arxiv.org/abs/2312.10997
2024 arXiv
-
[18]
Shailja Gupta, Rajesh Ranjan, and Surya Narayan Singh. 2024. Comprehensive Study on Sentiment Analysis: From Rule-based to modern LLM based system. arXiv:2409.09989 [cs.CL] https://arxiv.org/abs/2409.09989
2024 arXiv
-
[19]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations . https: //openreview.net/forum?id=nZeVKeeFYf9
2022
-
[20]
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bo- janowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised Dense Information Retrieval with Contrastive Learning. arXiv:2112.09118 [cs.IR] https://arxiv.org/abs/2112.09118
2022 arXiv
-
[21]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...
2023 arXiv
-
[22]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP...
2020 doi
-
[23]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019 doi
-
[24]
Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernon- court, Trung Bui, and Thien Huu Nguyen. 2023. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. In Findings of the Association for Computationa...
2023 doi
-
[25]
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv:2405.17428 [cs.CL] https: //arxiv.org/abs/2405.17428
2024 arXiv
-
[26]
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk
-
[27]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in ...
2020
-
[28]
Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them. Transactions of the Association for Computational Linguistics ...
2021
-
[29]
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards General Text Embeddings with Multi-stage Contrastive Learning. arXiv:2308.03281 [cs.CL] https://arxiv.org/abs/2308.03281
2023 arXiv
-
[30]
Jindřich Libovický, Rudolf Rosa, and Alexander Fraser. 2020. On the Language Neutrality of Pre-trained Multilingual Representations. In Findings of the Associ- ation for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computa...
2020 doi
-
[31]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning. Advances in neural information processing systems 36 (2024)
2024
-
[32]
Jiapeng Liu, Xiao Zhang, Dan Goldwasser, and Xiao Wang. 2020. Cross-Lingual Document Retrieval with Smooth Learning. InProceedings of the 28th International Conference on Computational Linguistics , Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on ...
2020 doi
-
[33]
Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Chankyu Lee, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Chatqa: Surpassing gpt-4 on conversational qa and rag. arXiv preprint arXiv:2401.10225 (2024)
2024 arXiv
-
[34]
Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye. 2024. Ovis: Structural Embedding Alignment for Multimodal Large Language Model. arXiv:2405.20797 [cs.CV] https://arxiv.org/abs/2405.20797
2024 arXiv
-
[35]
Kun Luo, Minghao Qin, Zheng Liu, Shitao Xiao, Jun Zhao, and Kang Liu. 2024. Large Language Models as Foundations for Next-Gen Dense Retrieval: A Com- prehensive Empirical Assessment. arXiv:2408.12194 [cs.CL] https://arxiv.org/ abs/2408.12194
2024 arXiv
-
[36]
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDer- mott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Financial Opinion Mining and Question Answering. In Companion Proceedings of the The Web Conference 2018 (Lyon, France) (WWW ’18)....
2018
-
[37]
Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, and Thien Huu Nguyen. 2024. ULLME: A Unified Framework for Large Language Model Embeddings with Generation-Augmented Learning. arXiv:2408.03402 [cs.CL] https://arxiv.org/ abs/2408.03402
2024 arXiv
-
[38]
Hieu Man and Thien Huu Nguyen. 2024. Counterfactual Augmentation for Robust Authorship Representation Learning. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24). Association for ...
2024
-
[39]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer Sentinel Mixture Models. In International Conference on Learning Repre- sentations. https://openreview.net/forum?id=Byj72udxe
2017
-
[40]
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. arXiv:1310.4546 [cs.CL] https://arxiv.org/abs/1310.4546
2013 arXiv
-
[41]
Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Aman- preet Singh, and Douwe Kiela. 2024. Generative Representational Instruction Tuning. arXiv:2402.09906 [cs.CL] https://arxiv.org/abs/2402.09906
2024 arXiv
-
[42]
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. MTEB: Massive Text Embedding Benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , Andreas Vlachos and Isabelle Augenstein (Eds.). Assoc...
2023 doi
-
[44]
Zhao, Yi Luan, Keith B
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2021. Large Dual Encoders Are Generalizable Retrievers. arXiv:2112.07899 [cs.IR] https://arxiv.org/abs/2112.07899
2021 arXiv
-
[45]
Hall, Daniel Cer, and Yinfei Yang
Jianmo Ni, Gustavo Hernández Ábrego, Noah Constant, Ji Ma, Keith B. Hall, Daniel Cer, and Yinfei Yang. 2021. Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models. arXiv:2108.08877 [cs.CL] https://arxiv.org/abs/ 2108.08877
2021 arXiv
-
[46]
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How Multilingual is Multi- lingual BERT?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Linguis...
2019 doi
-
[47]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceed- ings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Carreras (Eds.). A...
2016 doi
-
[48]
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Joban- putra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalak- shmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Deepak Kumar, Vivek Raghavan, Anoop Kunchukuttan, Pratyus...
2022 doi
-
[49]
Nils Reimers, Philip Beyer, and Iryna Gurevych. 2016. Task-Oriented Intrinsic Evaluation of Semantic Textual Similarity. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers , Yuji Matsumoto and Rashmi Prasad (Eds.). T...
2016
-
[50]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing (EMNLP-IJ...
2019 doi
-
[51]
Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y
Rafael A. Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, and Nicholas Andrews. 2021. Learning Univer- sal Authorship Representations. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing , ...
2021 doi
-
[52]
Stephen Robertson, Hugo Zaragoza, et al . 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends® in Information Retrieval 3, 4 (2009), 333–389
2009
-
[53]
Andrew Rosenberg and Julia Hirschberg. 2007. V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL) , ...
2007
-
[54]
Peng Shi, Rui Zhang, He Bai, and Jimmy Lin. 2021. Cross-Lingual Training of Dense Retrievers for Document Retrieval. In Proceedings of the 1st Workshop on Multilingual Representation Learning, Duygu Ataman, Alexandra Birch, Alexis Conneau, Orhan Firat, Sebastian Ruder, and Goz...
2021
-
[55]
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, and Han Xiao. 2024. jina-embeddings-v3: Multilingual Embeddings With Task LoRA. arXiv:2409.10173 [cs.CL] https://arxiv.org/ab...
2024 arXiv
-
[56]
Nandan Thakur, Jianmo Ni, Gustavo Hernández Ábrego, John Wieting, Jimmy Lin, and Daniel Cer. 2024. Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval. arXiv:2311.05800 [cs.IR] https://arxiv.org/abs/2311.05800
2024 arXiv
-
[57]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal
-
[58]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[59]
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the Best Counterargument without Prior Topic Knowledge. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Iryna Gurevych and Yusuke Miyao (Eds....
2018 doi
-
[60]
Boxin Wang, Wei Ping, Lawrence McAfee, Peng Xu, Bo Li, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Instructretro: Instruction tuning post retrieval- augmented pretraining. arXiv preprint arXiv:2310.07713 (2023)
2023 arXiv
-
[61]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2024. Text Embeddings by Weakly-Supervised Contrastive Pre-training. arXiv:2212.03533 [cs.CL] https://arxiv.org/abs/2212. 03533
2024 arXiv
-
[62]
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Improving Text Embeddings with Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre ...
2024 doi
-
[64]
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Multilingual E5 Text Embeddings: A Technical Report. arXiv preprint arXiv:2402.05672 (2024)
2024 arXiv
-
[65]
Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020. Extending Multilingual BERT to Low-Resource Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistic...
2020
-
[66]
Genta Indra Winata, Ruochen Zhang, and David Ifeoluwa Adelani. 2024. MINERS: Multilingual Language Models as Semantic Retrievers. arXiv:2406.07424 [cs.CL] https://arxiv.org/abs/2406.07424
2024 arXiv
-
[67]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[68]
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-Pack: Packaged Resources To Advance General Chinese Embedding. arXiv:2309.07597 [cs.CL]
2023 arXiv
-
[69]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...
2018 doi
-
[70]
Dongkeun Yoon, Joel Jang, Sungdong Kim, Seungone Kim, Sheikh Shafayat, and Minjoon Seo. 2024. LangBridge: Multilingual Reasoning Without Multilingual Supervision. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...
2024 doi
-
[71]
Bryan Zhang and Amita Misra. 2022. Machine translation impact in E-commerce multilingual search. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track , Yunyao Li and Angeliki Lazaridou (Eds.). Association for Computational L...
2022
-
[72]
Xin Zhang, Zehan Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, and Min Zhang. 2023. Language Models are Universal Embedders. arXiv preprint arXiv:2310.08232 (2023)
2023 arXiv
-
[73]
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. arXi...
2024 arXiv
-
[74]
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. PyTorch FSDP: Experiences...
2023 arXiv
-
[2018]
FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), Marilyn Walker, Heng Ji, and Amanda Ste...
2018 doi
-
[2020]
In Pro- ceedings of the 58th Annual Meeting of the Association for Computational Linguis- tics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.)
MLQA: Evaluating Cross-lingual Extractive Question Answering. In Pro- ceedings of the 58th Annual Meeting of the Association for Computational Linguis- tics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associa- tion for Computational Linguistics, Onl...
-
[2022]
In Ad- vances in Neural Information Processing Systems , S
Flamingo: a Visual Language Model for Few-Shot Learning. In Ad- vances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 23716–23736. https://proceedings.neurips.cc/paper_files...
2022
-
[2024]
arXiv:2402.03216 [cs.CL]
BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv:2402.03216 [cs.CL]
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.