Pith. sign in

MAIR: A Massive Benchmark for Evaluating Instructed Retrieval

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair.

citation-role summary

background 1

citation-polarity summary

fields

cs.IR 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

MIRB: Mathematical Information Retrieval Benchmark

cs.IR · 2025-05-21 · conditional · novelty 5.0

MIRB, a unified benchmark of four math retrieval tasks across 12 datasets, shows current retrieval models score far lower on premise retrieval than on semantic retrieval, and cross-encoder rerankers often hurt.

citing papers explorer

Showing 1 of 1 citing paper.

  • MIRB: Mathematical Information Retrieval Benchmark cs.IR · 2025-05-21 · conditional · none · ref 33 · internal anchor

    MIRB, a unified benchmark of four math retrieval tasks across 12 datasets, shows current retrieval models score far lower on premise retrieval than on semantic retrieval, and cross-encoder rerankers often hurt.