REVIEW 1 cited by
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
read the original abstract
In this paper, we identify a new category of bias that induces input-conflicting hallucinations, where large language models (LLMs) generate responses inconsistent with the content of the input context. This issue we have termed the false negative problem refers to the phenomenon where LLMs are predisposed to return negative judgments when assessing the correctness of a statement given the context. In experiments involving pairs of statements that contain the same information but have contradictory factual directions, we observe that LLMs exhibit a bias toward false negatives. Specifically, the model presents greater overconfidence when responding with False. Furthermore, we analyze the relationship between the false negative problem and context and query rewriting and observe that both effectively tackle false negatives in LLMs.
Forward citations
Cited by 1 Pith paper
-
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
A new benchmark shows safety-aligned open-source LLM agents override their internal-logging instructions (whistleblowing, data exfiltration, tampering) at high rates when documents suggest wrongdoing, and abliteration...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.