REVIEW 14 cited by
Arctic-Embed 2.0: Multilingual Retrieval Without Compromise
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents the training methodology of Arctic-Embed 2.0, a set of open-source text embedding models built for accurate and efficient multilingual retrieval. While prior works have suffered from degraded English retrieval quality, Arctic-Embed 2.0 delivers competitive retrieval quality on multilingual and English-only benchmarks, and supports Matryoshka Representation Learning (MRL) for efficient embedding storage with significantly lower compressed quality degradation compared to alternatives. We detail the design and implementation, presenting several important open research questions that arose during model development. We conduct experiments exploring these research questions and include extensive discussion aimed at fostering further discussion in this field.
Forward citations
Cited by 14 Pith papers
-
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
With matched open data and backbones, ColBERT-style late interaction turns English translate-train into multilingual generalization, while dense retrieval stays mostly inside the translated languages.
-
Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
Bekko a8m, with 7.7M active parameters, scores 56.2 on MMTEB Multilingual v2 Retrieval, beating mE5 models and BGE-M3, while a25m reaches 57.5, on par with gte-multilingual-base.
-
A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset
CUP, a 104-query Greek book retrieval benchmark, shows hybrid lexical-semantic retrieval outperforms both BM25 and dense-only methods in a real publisher catalog.
-
CORE-T: COherent REtrieval of Tables for Text-to-SQL
CORE-T uses LLM purpose metadata, a compatibility cache, and a single LLM call to select coherent joinable table sets, improving open-book multi-table text-to-SQL retrieval and execution accuracy.
-
LLMs Meet Isolation Kernel: Lightweight, Learning-free Binary Embeddings for Fast Retrieval
Isolation-kernel binary hashing (IKE) compresses LLM embeddings to a few hundred bytes per point with retrieval accuracy near the original and large speedups in exhaustive and ANN search.
-
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
CRUMB is a new benchmark for complex, multi-aspect retrieval tasks on which state-of-the-art retrieval models score poorly, and query rewriting does not rescue the best models.
-
Language Models Improve When Pretraining Data Matches Target Tasks
Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.
-
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
A new attack framework, TrojanClimb, shows that adversaries can place models with embedded backdoors or biases on public leaderboards while retaining competitive rankings, across text embeddings, text generation, spee...
-
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
JQL trains small multilingual quality scorers from LLM judgments and human annotations, and filtering pretraining data with them improves downstream multilingual model performance over heuristic baselines.
-
Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval
Amharic-specific dense retrieval models beat zero-shot multilingual baselines on a new headline-article benchmark, and a ColBERT variant achieves the top MRR@10 of 0.843.
-
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
By storing KV caches of prior reasoning traces and retrieving them during generation, LAG improves LLM agent accuracy and efficiency over standard agentic systems and reflection methods.
-
Training Sparse Mixture Of Experts Text Embedding Models
Nomic Embed v2 applies sparse mixture-of-experts upcycling to a multilingual biencoder, reporting competitive BEIR and MIRACL scores with fewer active parameters than dense models of similar size.
-
Generating consensus and dissent on massive discussion platforms with a semantic-vector model
A semantic-vector O(N) model on a 2D lattice generates consensus (β>0) or maximum dissent (β<0) by local copying of neighbor responses.
-
DoTA-RAG: Dynamic of Thought Aggregation RAG
DoTA-RAG combines query rewriting, namespace routing, dense retrieval, BM25 pruning, and reranking to answer questions over a 15M-document corpus, with reported correctness gains but fragile faithfulness under output caps.
Discussion (0). Continue with ORCID to comment.