Pith. sign in

REVIEW 5 major objections 5 minor 75 references

LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read LUSIFER shows that an English-only-trained connector can make an English-centric LLM embedding model competitive across 14 languages, beating E5-Mistral by 3.19 points.

desk verdict Solid English-only recipe for multilingual adaptation of LLM embedders, but the SOTA claim is oversold and the benchmark needs transparency. read the letter →

arxiv 2501.00874 v3 pith:HA2AZADK submitted 2025-01-01 cs.CL cs.IR

classification cs.CLcs.IR
keywords multilingualembeddingszero-shottransferlargelanguagemodelsdenseretrievaltextembeddingXLM-RMistral-7Bcontrastivelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

LUSIFER claims that an English-centric LLM embedding model can be made multilingual without seeing any multilingual training data, by routing input through a multilingual encoder and a small trainable connector. On the paper's new benchmark of 123 datasets across 14 languages, the method averages 62.63, beating the English-centric E5-Mistral baseline by 3.19 points and coming within two points of multilingual models trained on extensive multilingual data. The biggest gains are in medium- and low-resource languages, with Telugu improving by 22.15 points over E5-Mistral, and cross-lingual retrieval nearly doubling the strongest baseline on Indic languages. If the result holds, it means global text embeddings can be built from English-only resources and a minimal parameter budget.

What carries the argument

The load-bearing component is the connector: a two-layer feed-forward network that maps XLM-R's hidden states to Mistral's hidden dimension, plus one trainable token appended to the sequence, a pattern borrowed from multimodal alignment. Training has two stages, both on English data: stage one aligns encoder and connector with the frozen LLM using masked reconstruction and autoregressive completion; stage two finetunes everything end-to-end with contrastive learning, in-batch and hard negatives, bidirectional attention, and LoRA. The connector's job is to make the English-centric LLM interpret XLM-R's language-neutral vectors as if they were its own native text, so multilingual semantics transfer without multilingual supervision.

What would settle it

Train the identical connector and contrastive pipeline but replace XLM-R with a monolingual English encoder, then evaluate on the same 14-language benchmark; if the multilingual advantage over the English-only baseline disappears or reverses, the claimed transfer from a language-universal space is not what drives the gains.

Watch

Extended reading notes

Core claim

The paper's central claim is that the language-agnostic representations of XLM-R can be projected into the input space of an English-centric LLM (Mistral-7B) through a two-layer feed-forward connector plus a single trainable token, and that after two stages of English-only training the LLM produces strong embeddings for languages it rarely encountered during pretraining. The authors report state-of-the-art results on 10 of 14 languages in their benchmark, an average of 62.63 versus E5-Mistral's 59.44, and a cross-lingual average of 57.89 versus 52.14, including a near-doubling on the IndicCrosslingual dataset. They also show through t-SNE that LUSIFER's representations mix languages far more than E5-Mistral's, which they take as evidence of the intended language-agnostic behavior.

Load-bearing premise

The approach assumes that XLM-R's hidden representations are already language-agnostic enough that a connector trained only on English text can teach an English-centric LLM to correctly interpret them for any language.

Editorial extensions

If this is right

  • English-only training is enough to produce competitive multilingual embeddings across classification, clustering, reranking, retrieval, and STS, reducing the need for expensive multilingual corpora.
  • A small connector plus public English datasets can come within two points of large multilingual embedding models that required extensive multilingual training data.
  • Low-resource languages covered by XLM-R can see large embedding-quality gains, up to 22.15 points for Telugu over E5-Mistral.
  • Cross-lingual retrieval, where queries and documents are in different languages, improves by 5.75 points on average over the strongest English-centric baseline, with the largest gains on Indic languages.
  • The new 123-dataset, 14-language benchmark provides a common evaluation surface for future multilingual embedding research.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the connector recipe should transfer to other encoder-LLM pairs, so replacing XLM-R or Mistral with larger or instruction-tuned counterparts could scale the gains further.
  • Beyond the paper: because alignment is purely generative, the method inherits the multilingual coverage of the encoder; languages absent from XLM-R's pretraining would need encoder extensions before the connector can help.
  • Beyond the paper: task-level results in the paper show reranking lags the strongest baseline, so practitioners should check task-specific trade-offs before deploying the model in production multilingual search.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes LUSIFER, a two-stage method for adapting English-centric LLM embedding models to multilingual text embedding tasks without multilingual supervision for the embedding task itself. The architecture combines XLM-R-large as a multilingual encoder, a two-layer feed-forward connector with one trainable token, and Mistral-7B as the target LLM; the connector is first aligned with English masked-reconstruction and autoregressive objectives, and then all components are fine-tuned with contrastive learning on English retrieval data using LoRA. The authors evaluate on a newly assembled benchmark of 123 datasets across 14 languages and five embedding tasks, report improvements over the English-centric E5-Mistral model in 10 of 14 languages, report cross-lingual results on five datasets, and include ablations and a t-SNE visualization.

Significance. If the central claim were fully supported, the paper would demonstrate a practical recipe: a small connector trained only on English can transfer multilingual understanding from XLM-R to a large English-centric LLM, yielding large gains in medium- and low-resource languages. The paper has several strengths: the architecture is simple and clearly described, the two-stage training pipeline and hyperparameters are specified in detail, English-only training data is enumerated, ablations isolate the contributions of the alignment and fine-tuning stages, and the authors provide a code and data link. However, the evaluation currently rests on a self-constructed benchmark whose composition is not fully disclosed, and the headline 'state-of-the-art in 10 out of 14 languages' claim is only valid relative to E5-Mistral, not against the multilingual baselines included in the same table. Because these issues affect the central empirical claim, they need to be resolved before the paper can be recommended for acceptance.

major comments (5)
  1. [Section 4.3, Table 1] The statement that LUSIFER 'achieves state-of-the-art performance in 10 out of 14 languages' is only true when the comparison is restricted to E5-Mistral. Against the full baseline set in Table 1, LUSIFER is the highest-scoring model in only one language (Indonesian); BGE-M3 outperforms LUSIFER in 11 of 14 languages and Multilingual-E5-large outperforms it in 11 of 14. The average score of 62.63 is also lower than BGE-M3 (64.80) and Multilingual-E5-large (64.54). The state-of-the-art claim as written is therefore unsupported and should be revised to refer specifically to English-centric LLM baselines, not to the general class of multilingual embedding models.
  2. [Section 4.1, Tables 1 and 7-16] The aggregate comparison is difficult to interpret because the benchmark averages macro over languages with highly unequal numbers of datasets. The paper does not provide a per-language dataset count table, and the appendix shows that some languages have very few datasets (e.g., Farsi has six, and Telugu and Swahili have very few). The reported 3.19-point gain over E5-Mistral is therefore heavily driven by low-resource languages where E5-Mistral is expected to be weak, which does not isolate the connector's transfer mechanism. The authors should release the full dataset-to-language mapping and report per-language and per-dataset counts, ideally with standard errors or a sensitivity analysis over benchmark composition.
  3. [Section 4.1, Table 2, Abstract] The abstract states that cross-lingual evaluation uses four datasets, while Section 4.1 and Table 2 list five: Belebele, MLQA, STS17, STS22, and IndicCrosslingual. Additionally, the exact assignment of datasets to the 14 benchmark languages is not provided in the paper or in the linked repository description. This makes the central benchmark claim non-reproducible and must be fixed before the empirical results can be assessed.
  4. [Abstract, Section 1, Section 4.2] The claim that LUSIFER works 'without requiring explicit multilingual training data' is overstated. The method relies on XLM-R-large, which is pretrained on multilingual corpora; what the training pipeline avoids is explicit multilingual supervision for the embedding task. This distinction matters because the transfer mechanism may derive from XLM-R's multilingual pretraining rather than from the connector itself. The wording should be corrected throughout, including the abstract, to 'without multilingual supervision for the embedding task.'
  5. [Section 4.7, Figure 6] The evidence for the central 'language-agnostic universal space' hypothesis is a qualitative t-SNE plot of 200 samples from SIB200. This does not directly test whether the connector transfers multilingual semantics to the target LLM, and it is not supported by a quantitative metric. The Frozen Multilingual Encoder ablation in Table 5 shows that fine-tuning XLM-R matters, but a more direct test, such as comparing LUSIFER against running the same contrastive fine-tuning directly on XLM-R embeddings or measuring language-identification accuracy of the output embeddings, would strengthen the mechanism claim.
minor comments (5)
  1. [Table 5] The text says the Frozen Multilingual Encoder variant scored 56.74, but Table 5 reports 58.74; also, the sentence 'showed a 18.45 point' is missing the word 'drop' and should be reworded.
  2. [Figures 3-5] The per-dataset labels in the task-specific comparison figures are very small and difficult to read; the figures should be reformatted or supplemented with a table of per-task averages.
  3. [Section 3.1, Eq. (1)] The aligned hidden state equation is not numbered and the masking mechanism for padding tokens is described only in prose; adding a formal definition would improve clarity.
  4. [Section 4.1] The statement that the benchmark contains 123 datasets should clarify whether this counts dataset-language instances or unique datasets; the sum 48+24+24+22+5=123 suggests the former, and this should be stated explicitly.
  5. [Table 3] The hyperparameter table does not specify which modules LoRA is applied to (for example, attention projections only or also feed-forward layers); this detail is needed for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: LUSIFER's multilingual gains are evaluated on external benchmarks and do not reduce to the paper's own definitions or fitted targets.

full rationale

LUSIFER's central claim is an empirical transfer result: a two-layer connector plus LoRA finetuning on English-only data yields multilingual embedding gains. The evaluation is entirely against pre-existing datasets and external baselines. Nothing in Section 3 defines the target metric in terms of the model's own parameters: Stage 1 loss (masked LM + next-token prediction) and Stage 2 contrastive loss are trained on English corpora and evaluated on held-out multilingual and cross-lingual datasets. No multilingual evaluation score enters the training objective, and no parameter is fitted to the multilingual test sets. The 'language-universal space' is presented as a hypothesis supported by external citations [30,46] and by t-SNE, not as a theorem that the results are derived from. The only self-citations are [24] for the high/medium/low-resource language grouping and [37] in a list of prior embedding work; neither supplies a conclusion that the paper's gains reduce to. The paper does contain a claim-calibration issue worth noting as correctness risk, not circularity: Section 4.3 says 'state-of-the-art performance in 10 out of 14 languages' while Table 1 shows BGE-M3 and Multilingual-E5-large outperform LUSIFER in 11 and 12 of the 14 languages respectively; this affects the strength of the SOTA claim, but it is an internal-consistency/benchmark-composition concern rather than a circular reduction. Accordingly, no circular step meets the evidence bar.

Assumptions & free parameters 5 free parameters · 4 assumptions · 2 invented entities

The central claim is empirical; it rests on the transfer hypothesis, the choice of model components, and the benchmark construction. The free parameters are typical training hyperparameters chosen without sensitivity analysis. The axioms include the language-agnostic assumption and the validity of the self-constructed benchmark. The invented entities are the trainable token and the conceptual 'universal space', neither of which has independent falsifiable evidence.

free parameters (5)
  • mask_ratio = 0.5
    Random mask ratio for masked reconstruction in Stage 1 (Table 3). Chosen by hand with no sensitivity analysis.
  • num_hard_negatives = 7
    Number of hard negatives per query in contrastive finetuning (Table 3). Chosen by hand, no ablation.
  • lora_rank = 16
    LoRA rank for finetuning in Stage 2 (Table 3). Standard choice, not justified for this problem.
  • lora_alpha = 32
    LoRA scaling parameter (Table 3). Paired with rank 16.
  • connector_layers = 2
    Two-layer feed-forward connector plus one trainable token (Section 3.1). Architectural choice without ablation.
assumptions (4)
  • domain assumption XLM-R representations are language-agnostic enough for zero-shot transfer to other languages
    Invoked in Section 1 and Section 3, citing [30,46]. This is the foundation for the whole method.
  • ad hoc to paper A connector trained only on English can align multilingual encoder outputs with the target LLM's space
    Core hypothesis of LUSIFER; no proof, only empirical results.
  • domain assumption The 123-dataset, 14-language benchmark is a valid and unbiased measure of multilingual embedding quality
    Section 4.1 constructs the benchmark; dataset selection and aggregation are done by the authors.
  • domain assumption Macro-averaging language scores gives a meaningful overall performance metric
    Table 1 averages across languages equally, hiding English degradation.
invented entities (2)
  • trainable token t
    purpose: Appended to aligned hidden states to give the LLM a learned summary token
    Introduced in Section 3.1; it is a learned parameter with no external validation.
  • language-agnostic universal space
    purpose: Conceptual explanation for why the alignment transfers multilingual understanding
    Proposed in Section 1 and supported only by a t-SNE visualization in Section 4.7, not by a quantitative falsifiable test.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models." pith.science (2026). https://pith.science/paper/HA2AZADK

@misc{pith2026250100874,
  author       = {Pith},
  title        = {Pith review of: LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HA2AZADK}},
  note         = {Machine review of arXiv:2501.00874}
}
read the original abstract

Recent advancements in large language models (LLMs) based embedding models have established new state-of-the-art benchmarks for text embedding tasks, particularly in dense vector-based retrieval. However, these models predominantly focus on English, leaving multilingual embedding capabilities largely unexplored. To address this limitation, we present LUSIFER, a novel zero-shot approach that adapts LLM-based embedding models for multilingual tasks without requiring multilingual supervision. LUSIFER's architecture combines a multilingual encoder, serving as a language-universal learner, with an LLM-based embedding model optimized for embedding-specific tasks. These components are seamlessly integrated through a minimal set of trainable parameters that act as a connector, effectively transferring the multilingual encoder's language understanding capabilities to the specialized embedding model. Additionally, to comprehensively evaluate multilingual embedding performance, we introduce a new benchmark encompassing 5 primary embedding tasks, 123 diverse datasets, and coverage across 14 languages. Extensive experimental results demonstrate that LUSIFER significantly enhances the multilingual performance across various embedding tasks, particularly for medium and low-resource languages, without requiring explicit multilingual training data.

Figures

Figures reproduced from arXiv: 2501.00874 by the authors.

Figure 1
Figure 1. Overview of LUSIFER. Left: Align a multilingual encoder with the target English-centric LLM only using English [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Overview of tasks and datasets in our benchmark. Crosslingual datasets are marked with a blue shade. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Performance comparison of LUSIFER and baseline models on Classification and Clustering tasks. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Performance comparison of LUSIFER and baseline [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Performance comparison of LUSIFER and baseline models on Retrieval and STS tasks. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: t-SNE representation of 200 randomly samples from [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 15 canonical work pages

  1. [1]

    Alabi, Yanke Mao, Haonan Gao, and Annie En-Shiun Lee

    David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, and Annie En-Shiun Lee. 2024. SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects. arXiv:2309.07445 [cs.CL] https://arxiv.org/abs/2309.07445

  2. [3]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikoł aj Bińko...

  3. [4]

    Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the Cross-lingual Transferability of Monolingual Representations. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 4623–4637. doi:10...

  4. [5]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https://arxiv.org/abs/1611.09268

  5. [6]

    Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele Benchmark: a Parallel Reading Compre- hension Dataset in 122 Language Variants. InProceedings of the 62nd Annual Meet- ing of the Association for Computational Linguistic...

  6. [7]

    Rachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta, Sriram Ganapa- thy, Abhishek Bapna, Prateek Jain, and Partha Talukdar. 2024. LLM Augmented LLMs: Expanding Capabilities through Composition. In The Twelfth Interna- tional Conference on Learning Representations . https://openreview.net/forum?id= jjA4O1vJRz

  7. [8]

    Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. arXiv:2404.05961 [cs.CL] https://arxiv.org/abs/ 2404.05961

  8. [9]

    Bowman, Gabor Angeli, Christopher Potts, and Christopher D

    Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Man- ning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Pro- cessing, Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association for Computational Linguistics, Lisbon, Por...

Show all 75 references
  1. [10]

    Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu

  2. [11]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil- laume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learn- ing at Scale. In Proceedings of the 58th Annual Me...

  3. [12]

    Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bow- man, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating Cross- lingual Sentence Representations. In Proceedings of the 2018 Conference on Em- pirical Methods in Natural Language Processing , E...

  4. [13]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  5. [14]

    2024.PyTorch Lightning

    William Falcon and The PyTorch Lightning team. 2024.PyTorch Lightning. doi:10. 5281/zenodo.13254264

  6. [15]

    Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan. 2021. Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021) , Anna SIGIR ’25, July 13–18, 2025, Padua, Italy Hieu Man, ...

  7. [16]

    Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih ...

  8. [17]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] https://arxiv.org/abs/2312.10997

  9. [18]

    Shailja Gupta, Rajesh Ranjan, and Surya Narayan Singh. 2024. Comprehensive Study on Sentiment Analysis: From Rule-based to modern LLM based system. arXiv:2409.09989 [cs.CL] https://arxiv.org/abs/2409.09989

  10. [19]

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations . https: //openreview.net/forum?id=nZeVKeeFYf9

  11. [20]

    Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bo- janowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised Dense Information Retrieval with Contrastive Learning. arXiv:2112.09118 [cs.IR] https://arxiv.org/abs/2112.09118

  12. [21]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...

  13. [22]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP...

  14. [23]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  15. [24]

    Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernon- court, Trung Bui, and Thien Huu Nguyen. 2023. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. In Findings of the Association for Computationa...

  16. [25]

    Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv:2405.17428 [cs.CL] https: //arxiv.org/abs/2405.17428

  17. [26]

    Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk

  18. [27]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in ...

  19. [28]

    Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them. Transactions of the Association for Computational Linguistics ...

  20. [29]

    Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards General Text Embeddings with Multi-stage Contrastive Learning. arXiv:2308.03281 [cs.CL] https://arxiv.org/abs/2308.03281

  21. [30]

    Jindřich Libovický, Rudolf Rosa, and Alexander Fraser. 2020. On the Language Neutrality of Pre-trained Multilingual Representations. In Findings of the Associ- ation for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computa...

  22. [31]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning. Advances in neural information processing systems 36 (2024)

  23. [32]

    Jiapeng Liu, Xiao Zhang, Dan Goldwasser, and Xiao Wang. 2020. Cross-Lingual Document Retrieval with Smooth Learning. InProceedings of the 28th International Conference on Computational Linguistics , Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on ...

  24. [33]

    Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Chankyu Lee, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Chatqa: Surpassing gpt-4 on conversational qa and rag. arXiv preprint arXiv:2401.10225 (2024)

  25. [34]

    Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye. 2024. Ovis: Structural Embedding Alignment for Multimodal Large Language Model. arXiv:2405.20797 [cs.CV] https://arxiv.org/abs/2405.20797

  26. [35]

    Kun Luo, Minghao Qin, Zheng Liu, Shitao Xiao, Jun Zhao, and Kang Liu. 2024. Large Language Models as Foundations for Next-Gen Dense Retrieval: A Com- prehensive Empirical Assessment. arXiv:2408.12194 [cs.CL] https://arxiv.org/ abs/2408.12194

  27. [36]

    Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDer- mott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Financial Opinion Mining and Question Answering. In Companion Proceedings of the The Web Conference 2018 (Lyon, France) (WWW ’18)....

  28. [37]

    Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, and Thien Huu Nguyen. 2024. ULLME: A Unified Framework for Large Language Model Embeddings with Generation-Augmented Learning. arXiv:2408.03402 [cs.CL] https://arxiv.org/ abs/2408.03402

  29. [38]

    Hieu Man and Thien Huu Nguyen. 2024. Counterfactual Augmentation for Robust Authorship Representation Learning. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24). Association for ...

  30. [39]

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer Sentinel Mixture Models. In International Conference on Learning Repre- sentations. https://openreview.net/forum?id=Byj72udxe

  31. [40]

    Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. arXiv:1310.4546 [cs.CL] https://arxiv.org/abs/1310.4546

  32. [41]

    Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Aman- preet Singh, and Douwe Kiela. 2024. Generative Representational Instruction Tuning. arXiv:2402.09906 [cs.CL] https://arxiv.org/abs/2402.09906

  33. [42]

    Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. MTEB: Massive Text Embedding Benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , Andreas Vlachos and Isabelle Augenstein (Eds.). Assoc...

  34. [44]

    Zhao, Yi Luan, Keith B

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2021. Large Dual Encoders Are Generalizable Retrievers. arXiv:2112.07899 [cs.IR] https://arxiv.org/abs/2112.07899

  35. [45]

    Hall, Daniel Cer, and Yinfei Yang

    Jianmo Ni, Gustavo Hernández Ábrego, Noah Constant, Ji Ma, Keith B. Hall, Daniel Cer, and Yinfei Yang. 2021. Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models. arXiv:2108.08877 [cs.CL] https://arxiv.org/abs/ 2108.08877

  36. [46]

    Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How Multilingual is Multi- lingual BERT?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Linguis...

  37. [47]

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceed- ings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Carreras (Eds.). A...

  38. [48]

    Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Joban- putra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalak- shmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Deepak Kumar, Vivek Raghavan, Anoop Kunchukuttan, Pratyus...

  39. [49]

    Nils Reimers, Philip Beyer, and Iryna Gurevych. 2016. Task-Oriented Intrinsic Evaluation of Semantic Textual Similarity. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers , Yuji Matsumoto and Rashmi Prasad (Eds.). T...

  40. [50]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing (EMNLP-IJ...

  41. [51]

    Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y

    Rafael A. Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, and Nicholas Andrews. 2021. Learning Univer- sal Authorship Representations. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing , ...

  42. [52]

    Stephen Robertson, Hugo Zaragoza, et al . 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends® in Information Retrieval 3, 4 (2009), 333–389

  43. [53]

    Andrew Rosenberg and Julia Hirschberg. 2007. V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL) , ...

  44. [54]

    Peng Shi, Rui Zhang, He Bai, and Jimmy Lin. 2021. Cross-Lingual Training of Dense Retrievers for Document Retrieval. In Proceedings of the 1st Workshop on Multilingual Representation Learning, Duygu Ataman, Alexandra Birch, Alexis Conneau, Orhan Firat, Sebastian Ruder, and Goz...

  45. [55]

    Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, and Han Xiao. 2024. jina-embeddings-v3: Multilingual Embeddings With Task LoRA. arXiv:2409.10173 [cs.CL] https://arxiv.org/ab...

  46. [56]

    Nandan Thakur, Jianmo Ni, Gustavo Hernández Ábrego, John Wieting, Jimmy Lin, and Daniel Cer. 2024. Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval. arXiv:2311.05800 [cs.IR] https://arxiv.org/abs/2311.05800

  47. [57]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal

  48. [58]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...

  49. [59]

    Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the Best Counterargument without Prior Topic Knowledge. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Iryna Gurevych and Yusuke Miyao (Eds....

  50. [60]

    Boxin Wang, Wei Ping, Lawrence McAfee, Peng Xu, Bo Li, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Instructretro: Instruction tuning post retrieval- augmented pretraining. arXiv preprint arXiv:2310.07713 (2023)

  51. [61]

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2024. Text Embeddings by Weakly-Supervised Contrastive Pre-training. arXiv:2212.03533 [cs.CL] https://arxiv.org/abs/2212. 03533

  52. [62]

    Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Improving Text Embeddings with Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Lun-Wei Ku, Andre ...

  53. [64]

    Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Multilingual E5 Text Embeddings: A Technical Report. arXiv preprint arXiv:2402.05672 (2024)

  54. [65]

    Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020. Extending Multilingual BERT to Low-Resource Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistic...

  55. [66]

    Genta Indra Winata, Ruochen Zhang, and David Ifeoluwa Adelani. 2024. MINERS: Multilingual Language Models as Semantic Retrievers. arXiv:2406.07424 [cs.CL] https://arxiv.org/abs/2406.07424

  56. [67]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

  57. [68]

    Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-Pack: Packaged Resources To Advance General Chinese Embedding. arXiv:2309.07597 [cs.CL]

  58. [69]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...

  59. [70]

    Dongkeun Yoon, Joel Jang, Sungdong Kim, Seungone Kim, Sheikh Shafayat, and Minjoon Seo. 2024. LangBridge: Multilingual Reasoning Without Multilingual Supervision. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...

  60. [71]

    Bryan Zhang and Amita Misra. 2022. Machine translation impact in E-commerce multilingual search. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track , Yunyao Li and Angeliki Lazaridou (Eds.). Association for Computational L...

  61. [72]

    Xin Zhang, Zehan Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, and Min Zhang. 2023. Language Models are Universal Embedders. arXiv preprint arXiv:2310.08232 (2023)

  62. [73]

    Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. arXi...

  63. [74]

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. PyTorch FSDP: Experiences...

  64. [2018]

    FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), Marilyn Walker, Heng Ji, and Amanda Ste...

  65. [2020]

    In Pro- ceedings of the 58th Annual Meeting of the Association for Computational Linguis- tics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.)

    MLQA: Evaluating Cross-lingual Extractive Question Answering. In Pro- ceedings of the 58th Annual Meeting of the Association for Computational Linguis- tics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associa- tion for Computational Linguistics, Onl...

  66. [2022]

    In Ad- vances in Neural Information Processing Systems , S

    Flamingo: a Visual Language Model for Few-Shot Learning. In Ad- vances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 23716–23736. https://proceedings.neurips.cc/paper_files...

  67. [2024]

    arXiv:2402.03216 [cs.CL]

    BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv:2402.03216 [cs.CL]

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.