D e C o R e: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations

Gema, Aryo Pradipta, Jin, Chen, Abdulaal, Ahmed, Diethe, Tom, Teare, Philip Alexander, Alex, Beatrice · 2025 · DOI 10.18653/v1/2025.findings-emnlp.531

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open at publisher browse 1 citing papers

representative citing papers

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

cs.CL · 2026-07-01 · unverdicted · novelty 7.0

LOCOS scores attention heads via OV-circuit output projection onto answer-token unembedding directions and identifies non-literal retrieval heads whose ablation collapses performance on non-literal benchmarks more than prior literal-copy detectors.

citing papers explorer

Showing 1 of 1 citing paper.

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads cs.CL · 2026-07-01 · unverdicted · none · ref 21
LOCOS scores attention heads via OV-circuit output projection onto answer-token unembedding directions and identifies non-literal retrieval heads whose ablation collapses performance on non-literal benchmarks more than prior literal-copy detectors.

D e C o R e: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations

fields

years

verdicts

representative citing papers

citing papers explorer