REVIEW 10 cited by
mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The MS MARCO ranking dataset has been widely used for training deep learning models for IR tasks, achieving considerable effectiveness on diverse zero-shot scenarios. However, this type of resource is scarce in languages other than English. In this work, we present mMARCO, a multilingual version of the MS MARCO passage ranking dataset comprising 13 languages that was created using machine translation. We evaluated mMARCO by finetuning monolingual and multilingual reranking models, as well as a multilingual dense retrieval model on this dataset. We also evaluated models finetuned using the mMARCO dataset in a zero-shot scenario on Mr. TyDi dataset, demonstrating that multilingual models finetuned on our translated dataset achieve superior effectiveness to models finetuned on the original English version alone. Our experiments also show that a distilled multilingual reranker is competitive with non-distilled models while having 5.4 times fewer parameters. Lastly, we show a positive correlation between translation quality and retrieval effectiveness, providing evidence that improvements in translation methods might lead to improvements in multilingual information retrieval. The translated datasets and finetuned models are available at https://github.com/unicamp-dl/mMARCO.
Forward citations
Cited by 10 Pith papers
-
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
With matched open data and backbones, ColBERT-style late interaction turns English translate-train into multilingual generalization, while dense retrieval stays mostly inside the translated languages.
-
Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings
Dense retrievers frequently miss simple fine-grained entity and event matches in image captions, and keyword-based training fixes this on the new CapRetrieval benchmark but can degrade overall semantic retrieval.
-
Disentangled Contrastive Learning for Zero-Shot Multilingual Dense Retrieval
A disentangled semantic/linguistic subspace training method improves zero-shot multilingual dense retrieval on mMARCO and MIRACL without target-language retrieval labels.
-
LAMAR: An Open Language-Aware Multilingual Alignment Reranker
A 0.6B multilingual reranker trained to prefer query-language documents over semantically equivalent other-language documents beats existing rerankers on language-coherence tests without collapsing general reranking quality.
-
A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset
CUP, a 104-query Greek book retrieval benchmark, shows hybrid lexical-semantic retrieval outperforms both BM25 and dense-only methods in a real publisher catalog.
-
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
KaLM-Reranker-V1 uses encoder–decoder FBNL with Matryoshka pooling to match Qwen3-class reranking quality at substantially lower online cost.
-
HyReC: Exploring Hybrid-based Retriever for Chinese
HyReC unifies dense, lexicon, and learned word-segment retrieval into one model and reports improved C-MTEB retrieval scores for Chinese.
-
Unify Graph Learning with Text: Unleashing LLM Potentials for Session Search
A session graph serialized into symbolic text, plus self-supervised graph pre-training tasks, lets an LLM outperform existing session search rankers on AOL and Tiangong-ST.
-
Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
Semantically parallel queries in 24 European languages get inconsistent rankings from BM25 and neural retrievers; a KL-divergence alignment loss (LaKDA) reduces the inconsistency.
-
LAGO: Few-shot Crosslingual Embedding Inversion Attacks via Language Similarity-Aware Graph Optimization
LAGO shows that constraining alignment matrices of linguistically similar languages to be close improves few-shot cross-lingual embedding inversion accuracy over independent per-language baselines.
Discussion (0). Continue with ORCID to comment.