Pith. sign in

REVIEW 3 cited by

Simple Entity-Centric Questions Challenge Dense Retrievers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.08535 v3 pith:4WSLLAIH submitted 2021-09-17 cs.CL cs.IR

classification cs.CLcs.IR
keywords densequestionmodelsretrieverssimpledemonstratefirstonly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-domain question answering has exploded in popularity recently due to the success of dense retrieval models, which have surpassed sparse models using only a few supervised training examples. However, in this paper, we demonstrate current dense models are not yet the holy grail of retrieval. We first construct EntityQuestions, a set of simple, entity-rich questions based on facts from Wikidata (e.g., "Where was Arve Furset born?"), and observe that dense retrievers drastically underperform sparse methods. We investigate this issue and uncover that dense retrievers can only generalize to common entities unless the question pattern is explicitly observed during training. We discuss two simple solutions towards addressing this critical problem. First, we demonstrate that data augmentation is unable to fix the generalization problem. Second, we argue a more robust passage encoder helps facilitate better question adaptation using specialized question encoders. We hope our work can shed light on the challenges in creating a robust, universal dense retriever that works well across different input distributions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment

    cs.IR 2025-07 conditional novelty 6.0 of 10

    AQE uses retrieval rank as a preference signal to fine-tune T0 with RSFT and DPO, beating generate-then-filter baselines on four QA datasets.

  2. From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Fine-tuning LLMs on known versus unknown facts creates a factuality gap that in-context prompting can largely erase, according to experiments and a knowledge-graph model.

  3. CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment

    cs.IR 2025-02 conditional novelty 6.0 of 10

    CliniQ is a public EHR retrieval benchmark with 77,206 LLM-annotated relevance judgments, showing that BM25 is a strong baseline and that semantic matches drive dense-retriever gains.

Pith tools