Pith. sign in

REVIEW 10 cited by

Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05374 v1 pith:6LKHVFJ7 submitted 2024-05-08 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords modelsembeddingmodelarctic-embedreciperetrievaltexttraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report describes the training dataset creation and recipe behind the family of \texttt{arctic-embed} text embedding models (a set of five models ranging from 22 to 334 million parameters with weights open-sourced under an Apache-2 license). At the time of their release, each model achieved state-of-the-art retrieval accuracy for models of their size on the MTEB Retrieval leaderboard, with the largest model, arctic-embed-l outperforming closed source embedding models such as Cohere's embed-v3 and Open AI's text-embed-3-large. In addition to the details of our training recipe, we have provided several informative ablation studies, which we believe are the cause of our model performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

    cs.CL 2026-07 conditional novelty 7.0 of 10

    With matched open data and backbones, ColBERT-style late interaction turns English translate-train into multilingual generalization, while dense retrieval stays mostly inside the translated languages.

  2. FlexOlmo: Open Language Models for Flexible Data Use

    cs.CL 2025-07 conditional novelty 7.0 of 10

    FlexOlmo merges independently trained language-model experts, trained on private data, into a single mixture-of-experts model without joint training.

  3. Finding the Right Tables and Columns: A Benchmark and Corpus-Adaptive Embeddings for SQL Schema Retrieval

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Schema retrieval can be benchmarked as retrieval, and corpus-adaptive fine-tuning lifts a 305M embedder to 75.6 recall@10, rivaling 4–8B models.

  4. Removing Noise, not Finding Gold: Quality Filtering for Large-Scale Pretraining

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Classifier-based quality filtering for LLM pretraining improves downstream tasks by implicitly filtering the reference high-quality set rather than by mimicking it, and its quality scores fail a data-conditioning test.

  5. Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI

    cs.DC 2025-07 conditional novelty 6.0 of 10

    Arctic Inference introduces Shift Parallelism, dynamic switching between tensor and sequence parallelism, achieving faster LLM inference and higher embedding throughput in a single deployment.

  6. Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A new attack framework, TrojanClimb, shows that adversaries can place models with embedded backdoors or biases on public leaderboards while retaining competitive rankings, across text embeddings, text generation, spee...

  7. MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MedVideoCap-55K, a 55,803-clip caption-rich medical video dataset, enables MedGen, a LoRA fine-tune of HunyuanVideo that reports top open-source scores and near-commercial quality on medical video benchmarks.

  8. DeRAG: Black-box Adversarial Attacks on Multiple Retrieval-Augmented Generation Applications via Prompt Injection

    cs.AI 2025-07 conditional novelty 5.0 of 10

    DeRAG shows that five or fewer tokens found by differential evolution can make black-box RAG retrievers rank a chosen wrong document near the top on small BEIR subsets.

  9. Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data

    cs.IR 2025-05 conditional novelty 5.0 of 10

    Contrastive fine-tuning often degrades strong dense retrievers, while combining cross-encoder listwise distillation with diverse synthetic queries consistently improves them.

  10. Boosting Data Utilization for Multilingual Dense Retrieval

    cs.IR 2025-09 conditional novelty 4.0 of 10

    A three-stage data-utilization pipeline for multilingual dense retrieval, combining ensemble hard-negative mining, LLM-based filtering/generation, and monolingual topic-diverse mini-batches, improves MIRACL nDCG@10 by...

Pith tools