Pith. sign in

REVIEW 7 cited by

INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.14334 v1 pith:ZC3VRHPU submitted 2024-02-22 cs.CL

classification cs.CL
keywords searchinformationinstructionsretrievalretrieversusersabilitybenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the critical need to align search targets with users' intention, retrievers often only prioritize query information without delving into the users' intended search context. Enhancing the capability of retrievers to understand intentions and preferences of users, akin to language model instructions, has the potential to yield more aligned search targets. Prior studies restrict the application of instructions in information retrieval to a task description format, neglecting the broader context of diverse and evolving search scenarios. Furthermore, the prevailing benchmarks utilized for evaluation lack explicit tailoring to assess instruction-following ability, thereby hindering progress in this field. In response to these limitations, we propose a novel benchmark,INSTRUCTIR, specifically designed to evaluate instruction-following ability in information retrieval tasks. Our approach focuses on user-aligned instructions tailored to each query instance, reflecting the diverse characteristics inherent in real-world search scenarios. Through experimental analysis, we observe that retrievers fine-tuned to follow task-style instructions, such as INSTRUCTOR, can underperform compared to their non-instruction-tuned counterparts. This underscores potential overfitting issues inherent in constructing retrievers trained on existing instruction-aware retrieval datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EXCISE: Query-Side Exclusion for Late-Interaction Retrieval

    cs.IR 2026-08 conditional novelty 6.0 of 10

    EXCISE adds two small query-side modules and a demotion rule to frozen ColBERT indexes, raising exclusion success@10 on ExcluIR from 0.058 to 0.691 without re-indexing and without losing ordinary retrieval quality on ...

  2. UEmbed: Unified Sparse and Dense Multimodal Embeddings

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UEmbed uses 16 special tokens over a partitioned vocabulary to make a decoder-only multimodal model emit dense and sparse embeddings in one forward pass; the 9B model scores 71.8 dense / 71.0 sparse on MMEB-v2.

  3. Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new 1,012-question benchmark shows LLMs often fail instructions that deliberately invert common training conventions, revealing a measurable gap in counterintuitive instruction following.

  4. Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    GSTransform trains a shallow linear transformation on a few thousand LLM-labeled examples to make precomputed text embeddings adapt to user instructions in real time.

  5. LLMs can be easily Confused by Instructional Distractions

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new benchmark, DIM-Bench, shows that LLMs frequently follow instructions hidden inside the target input rather than the user's actual instruction, even when explicitly told to ignore them.

  6. Beyond Sequential Reranking: Reranker-Guided Search Improves Reasoning Intensive Retrieval

    cs.IR 2025-09 conditional novelty 5.0 of 10

    Reranker-Guided-Search, a greedy graph search steered by reranker scores, outperforms sequential top-k reranking under a fixed budget on three reasoning-intensive retrieval benchmarks.

  7. Towards Better Instruction Following Retrieval Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A new training corpus and embedding model improve instruction-following p-MRR by up to 9 points on FollowIR, MAIR, and Bright benchmarks.

Pith tools