Pith. sign in

REVIEW 5 cited by

Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.02802 v1 pith:BUGTWWTS submitted 2024-12-03 cs.AI

classification cs.AI
keywords modelbehaviorlanguagesycophantictrustlargeparticipantsuser
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sycophancy refers to the tendency of a large language model to align its outputs with the user's perceived preferences, beliefs, or opinions, in order to look favorable, regardless of whether those statements are factually correct. This behavior can lead to undesirable consequences, such as reinforcing discriminatory biases or amplifying misinformation. Given that sycophancy is often linked to human feedback training mechanisms, this study explores whether sycophantic tendencies negatively impact user trust in large language models or, conversely, whether users consider such behavior as favorable. To investigate this, we instructed one group of participants to answer ground-truth questions with the assistance of a GPT specifically designed to provide sycophantic responses, while another group used the standard version of ChatGPT. Initially, participants were required to use the language model, after which they were given the option to continue using it if they found it trustworthy and useful. Trust was measured through both demonstrated actions and self-reported perceptions. The findings consistently show that participants exposed to sycophantic behavior reported and exhibited lower levels of trust compared to those who interacted with the standard version of the model, despite the opportunity to verify the accuracy of the model's output.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

    cs.CL 2026-07 conditional novelty 6.0 of 10

    On 216 news headlines about three conflicts, GPT-5.2's sympathy judgments correlate 0.79 with a representative UK panel, Mistral's only 0.41, with significant demographic variation.

  2. Evaluating Intra-firm LLM Alignment Strategies in Business Contexts

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Firms should intentionally align AI assistants' embedded perspectives using supportive, adversarial, or diverse strategies to protect workplace culture and moral norms.

  3. Effects of Personality- and Opinion-Alignment in Human-AI Interaction

    cs.HC 2025-11 conditional novelty 5.0 of 10

    People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

  4. Exploring and Mitigating Fawning Hallucinations in Large Language Models

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A contrastive decoding method that contrasts a misleading prompt against a neutral rewrite reduces fawning hallucinations in LLMs, though most of the gain comes from the neutral prompt itself.

  5. Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages

    cs.CL 2025-06 conditional novelty 3.0 of 10

    Relabeling hate speech as metaphor pairs (red/green, summer/winter) in prompts raises Llama2's F1 on a 500-item Bengali subsample to 95.89, though the gain is reported without matched test-set comparisons or error bars.

Pith tools