REVIEW 7 cited by
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The interactive nature of Large Language Models (LLMs) theoretically allows models to refine and improve their answers, yet systematic analysis of the multi-turn behavior of LLMs remains limited. In this paper, we propose the FlipFlop experiment: in the first round of the conversation, an LLM completes a classification task. In a second round, the LLM is challenged with a follow-up phrase like "Are you sure?", offering an opportunity for the model to reflect on its initial answer, and decide whether to confirm or flip its answer. A systematic study of ten LLMs on seven classification tasks reveals that models flip their answers on average 46% of the time and that all models see a deterioration of accuracy between their first and final prediction, with an average drop of 17% (the FlipFlop effect). We conduct finetuning experiments on an open-source LLM and find that finetuning on synthetically created data can mitigate - reducing performance deterioration by 60% - but not resolve sycophantic behavior entirely. The FlipFlop experiment illustrates the universality of sycophantic behavior in LLMs and provides a robust framework to analyze model behavior and evaluate future models.
Forward citations
Cited by 7 Pith papers
-
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
Across six open-weight LLMs and seven datasets, a speaker-free wrong-answer assertion alone flips 66.5% of initially correct answers, versus 10.3% for a plain re-ask; source labels mainly add a modest increment above ...
-
MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs
Across 600 five-turn medical dialogues, most of 20 LLMs shift from safe stances to unsafe agreement once patients apply escalating pressure.
-
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
Showing LLM agents precomputed rankings of their peers' sycophancy improves multi-agent discussion accuracy by ~10.5 absolute points and reduces agreement with incorrect user stances.
-
SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care
In simulated emergency-department encounters, 20 LLMs acquiesced to patient pressure for unindicated CT scans, antibiotics, or opioids at rates varying from 0% to 100%, with no clear link to model capability or recency.
-
Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems
LLM agents flip their answers more when their own confidence is low and their peer seems confident, and the format and order of peer information can amplify or dampen this herd behavior.
-
B-score: Detecting biases in large language models using response history
LLMs self-correct toward uniform answers in multi-turn repetition, and the gap between single-turn and multi-turn answer rates (B-score) flags biased answers better than verbalized confidence.
-
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.
Discussion (0). Continue with ORCID to comment.