Pith. sign in

REVIEW 7 cited by

When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09410 v4 pith:ERZOW3RI submitted 2023-11-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelslanguagelargebehaviourllmssycophanticuserswhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models have been demonstrating broadly satisfactory generative abilities for users, which seems to be due to the intensive use of human feedback that refines responses. Nevertheless, suggestibility inherited via human feedback improves the inclination to produce answers corresponding to users' viewpoints. This behaviour is known as sycophancy and depicts the tendency of LLMs to generate misleading responses as long as they align with humans. This phenomenon induces bias and reduces the robustness and, consequently, the reliability of these models. In this paper, we study the suggestibility of Large Language Models (LLMs) to sycophantic behaviour, analysing these tendencies via systematic human-interventions prompts over different tasks. Our investigation demonstrates that LLMs have sycophantic tendencies when answering queries that involve subjective opinions and statements that should elicit a contrary response based on facts. In contrast, when faced with math tasks or queries with an objective answer, they, at various scales, do not follow the users' hints by demonstrating confidence in generating the correct answers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Open-ended LLM conformity is not captured by answer flips: wrong peers degrade revisions, and judges' ratings shift when peer context is visible, so evaluation must be modeled explicitly.

  2. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  3. The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs

    cs.HC 2026-04 conditional novelty 6.0 of 10

    Sycophancy is persona-conditional: a strongly-aligned model stays within 5pp across personas while a lightly-aligned one spans 45pp, so persona safety requires per-model auditing.

  4. DocPrism: Multi-lingual Detection of Incorrectness Inconsistencies between Code and Documentation

    cs.SE 2025-10 conditional novelty 6.0 of 10

    A zero-shot LLM prompting scheme (local categorization + external filtering) detects code-documentation incorrectness with low flag rates and about 0.6 precision across Python, TypeScript, C++, and Java.

  5. TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Token-level difference-weighted preference optimization on minimal-edit pairs reduces sycophancy in autism-intervention LLMs while preserving intervention skill.

  6. The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis

    cs.AI 2025-09 conditional novelty 4.0 of 10

    Six large language models consistently rated care and virtue outcomes as most moral and libertarian outcomes as least moral across 54 AI-generated dilemma variants, with reasoning models more context-sensitive but les...

  7. Self-Critique-Guided Curiosity Refinement: Enhancing Honesty and Helpfulness in Large Language Models via In-Context Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Adding a self-critique and refinement step to curiosity-driven prompting improves GPT-4o-judged honesty and helpfulness scores on HONESET by 1.4% to 4.3% across ten LLMs.

Pith tools