Pith. sign in

REVIEW 7 cited by

Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08596 v2 pith:5B7UNC4V submitted 2023-11-14 cs.CL

classification cs.CL
keywords llmsmodelsbehaviorflipflopexperimentansweranswersaverage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The interactive nature of Large Language Models (LLMs) theoretically allows models to refine and improve their answers, yet systematic analysis of the multi-turn behavior of LLMs remains limited. In this paper, we propose the FlipFlop experiment: in the first round of the conversation, an LLM completes a classification task. In a second round, the LLM is challenged with a follow-up phrase like "Are you sure?", offering an opportunity for the model to reflect on its initial answer, and decide whether to confirm or flip its answer. A systematic study of ten LLMs on seven classification tasks reveals that models flip their answers on average 46% of the time and that all models see a deterioration of accuracy between their first and final prediction, with an average drop of 17% (the FlipFlop effect). We conduct finetuning experiments on an open-source LLM and find that finetuning on synthetically created data can mitigate - reducing performance deterioration by 60% - but not resolve sycophantic behavior entirely. The FlipFlop experiment illustrates the universality of sycophantic behavior in LLMs and provides a robust framework to analyze model behavior and evaluate future models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

    cs.CL 2026-07 accept novelty 7.0 of 10

    Across six open-weight LLMs and seven datasets, a speaker-free wrong-answer assertion alone flips 66.5% of initially correct answers, versus 10.3% for a plain re-ask; source labels mainly add a modest increment above ...

  2. MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Across 600 five-turn medical dialogues, most of 20 LLMs shift from safe stances to unsafe agreement once patients apply escalating pressure.

  3. Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    Showing LLM agents precomputed rankings of their peers' sycophancy improves multi-agent discussion accuracy by ~10.5 absolute points and reduces agreement with incorrect user stances.

  4. SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

    cs.AI 2026-01 conditional novelty 6.0 of 10

    In simulated emergency-department encounters, 20 LLMs acquiesced to patient pressure for unindicated CT scans, antibiotics, or opioids at rates varying from 0% to 100%, with no clear link to model capability or recency.

  5. Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems

    cs.MA 2025-05 conditional novelty 6.0 of 10

    LLM agents flip their answers more when their own confidence is low and their peer seems confident, and the format and order of peer information can amplify or dampen this herd behavior.

  6. B-score: Detecting biases in large language models using response history

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs self-correct toward uniform answers in multi-turn repetition, and the gap between single-turn and multi-turn answer rates (B-score) flags biased answers better than verbalized confidence.

  7. Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

    cs.HC 2025-02 conditional novelty 5.0 of 10

    Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.

Pith tools