LoHoSearch is a new benchmark of 544 KG-constructed questions across 11 domains where the strongest search agent scores 34.74% and context strategies add at most 6.8%.
Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation
4 Pith papers cite this work, alongside 14 external citations. Polarity classification is still indexing.
years
2026 4representative citing papers
LiveBrowseComp shows search agents rely on intrinsic knowledge on standard benchmarks, with scores dropping 25-40 points and closed-book accuracy below 2% on questions about facts from the prior 90 days.
Counterfactual no-search vs forced-search outcomes yield a model-specific oracle that trains search-routing policies, raising macro-F1 from ~0.71 to ~0.82–0.84 on oracle-eligible examples.
CERA fine-tunes a dense retriever with triplet contrastive learning plus attention alignment to human rationales, claiming better retrieval effectiveness and faithfulness on clinical trial reports than Contriever and standard hard-negative baselines.
citing papers explorer
-
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling
LoHoSearch is a new benchmark of 544 KG-constructed questions across 11 domains where the strongest search agent scores 34.74% and context strategies add at most 6.8%.
-
LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?
LiveBrowseComp shows search agents rely on intrinsic knowledge on standard benchmarks, with scores dropping 25-40 points and closed-book accuracy below 2% on questions about facts from the prior 90 days.
-
When Should LLMs Search? Counterfactual Supervision for Search Routing
Counterfactual no-search vs forced-search outcomes yield a model-specific oracle that trains search-routing policies, raising macro-F1 from ~0.71 to ~0.82–0.84 on oracle-eligible examples.
-
Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG
CERA fine-tunes a dense retriever with triplet contrastive learning plus attention alignment to human rationales, claiming better retrieval effectiveness and faithfulness on clinical trial reports than Contriever and standard hard-negative baselines.