Pith. sign in

REVIEW 14 cited by

Arctic-Embed 2.0: Multilingual Retrieval Without Compromise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04506 v2 pith:LRVJHGFP submitted 2024-12-03 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords retrievalarctic-embedmultilingualqualitydiscussionefficientembeddingquestions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents the training methodology of Arctic-Embed 2.0, a set of open-source text embedding models built for accurate and efficient multilingual retrieval. While prior works have suffered from degraded English retrieval quality, Arctic-Embed 2.0 delivers competitive retrieval quality on multilingual and English-only benchmarks, and supports Matryoshka Representation Learning (MRL) for efficient embedding storage with significantly lower compressed quality degradation compared to alternatives. We detail the design and implementation, presenting several important open research questions that arose during model development. We conduct experiments exploring these research questions and include extensive discussion aimed at fostering further discussion in this field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    cs.CL 2026-07 conditional novelty 7.0 of 10

    With matched open data and backbones, ColBERT-style late interaction turns English translate-train into multilingual generalization, while dense retrieval stays mostly inside the translated languages.

  2. Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders

    cs.IR 2026-07 conditional novelty 7.0 of 10

    Bekko a8m, with 7.7M active parameters, scores 56.2 on MMTEB Multilingual v2 Retrieval, beating mE5 models and BGE-M3, while a25m reaches 57.5, on par with gte-multilingual-base.

  3. A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CUP, a 104-query Greek book retrieval benchmark, shows hybrid lexical-semantic retrieval outperforms both BM25 and dense-only methods in a real publisher catalog.

  4. CORE-T: COherent REtrieval of Tables for Text-to-SQL

    cs.CL 2026-01 conditional novelty 6.0 of 10

    CORE-T uses LLM purpose metadata, a compatibility cache, and a single LLM call to select coherent joinable table sets, improving open-book multi-table text-to-SQL retrieval and execution accuracy.

  5. LLMs Meet Isolation Kernel: Lightweight, Learning-free Binary Embeddings for Fast Retrieval

    cs.IR 2026-01 unverdicted novelty 6.0 of 10

    Isolation-kernel binary hashing (IKE) compresses LLM embeddings to a few hundred bytes per point with retrieval accuracy near the original and large speedups in exhaustive and ANN search.

  6. Benchmarking Information Retrieval Models on Complex Retrieval Tasks

    cs.IR 2025-09 conditional novelty 6.0 of 10

    CRUMB is a new benchmark for complex, multi-aspect retrieval tasks on which state-of-the-art retrieval models score poorly, and query rewriting does not rescue the best models.

  7. Language Models Improve When Pretraining Data Matches Target Tasks

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.

  8. Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A new attack framework, TrojanClimb, shows that adversaries can place models with embedded backdoors or biases on public leaderboards while retaining competitive rankings, across text embeddings, text generation, spee...

  9. Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    JQL trains small multilingual quality scorers from LLM judgments and human annotations, and filtering pretraining data with them improves downstream multilingual model performance over heuristic baselines.

  10. Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

    cs.IR 2025-05 conditional novelty 6.0 of 10

    Amharic-specific dense retrieval models beat zero-shot multilingual baselines on a new headline-article benchmark, and a ColBERT variant achieves the top MRR@10 of 0.843.

  11. Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    By storing KV caches of prior reasoning traces and retrieving them during generation, LAG improves LLM agent accuracy and efficiency over standard agentic systems and reflection methods.

  12. Training Sparse Mixture Of Experts Text Embedding Models

    cs.CL 2025-02 reject novelty 6.0 of 10

    Nomic Embed v2 applies sparse mixture-of-experts upcycling to a multilingual biencoder, reporting competitive BEIR and MIRACL scores with fewer active parameters than dense models of similar size.

  13. Generating consensus and dissent on massive discussion platforms with a semantic-vector model

    physics.soc-ph 2026-01 conditional novelty 5.0 of 10

    A semantic-vector O(N) model on a 2D lattice generates consensus (β>0) or maximum dissent (β<0) by local copying of neighbor responses.

  14. DoTA-RAG: Dynamic of Thought Aggregation RAG

    cs.CL 2025-06 conditional novelty 4.0 of 10

    DoTA-RAG combines query rewriting, namespace routing, dense retrieval, BM25 pruning, and reranking to answer questions over a 15M-document corpus, with reported correctness gains but fragile faithfulness under output caps.

Pith tools