REVIEW 5 cited by
Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sycophancy refers to the tendency of a large language model to align its outputs with the user's perceived preferences, beliefs, or opinions, in order to look favorable, regardless of whether those statements are factually correct. This behavior can lead to undesirable consequences, such as reinforcing discriminatory biases or amplifying misinformation. Given that sycophancy is often linked to human feedback training mechanisms, this study explores whether sycophantic tendencies negatively impact user trust in large language models or, conversely, whether users consider such behavior as favorable. To investigate this, we instructed one group of participants to answer ground-truth questions with the assistance of a GPT specifically designed to provide sycophantic responses, while another group used the standard version of ChatGPT. Initially, participants were required to use the language model, after which they were given the option to continue using it if they found it trustworthy and useful. Trust was measured through both demonstrated actions and self-reported perceptions. The findings consistently show that participants exposed to sycophantic behavior reported and exhibited lower levels of trust compared to those who interacted with the standard version of the model, despite the opportunity to verify the accuracy of the model's output.
Forward citations
Cited by 5 Pith papers
-
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
On 216 news headlines about three conflicts, GPT-5.2's sympathy judgments correlate 0.79 with a representative UK panel, Mistral's only 0.41, with significant demographic variation.
-
Evaluating Intra-firm LLM Alignment Strategies in Business Contexts
Firms should intentionally align AI assistants' embedded perspectives using supportive, adversarial, or diverse strategies to protect workplace culture and moral norms.
-
Effects of Personality- and Opinion-Alignment in Human-AI Interaction
People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.
-
Exploring and Mitigating Fawning Hallucinations in Large Language Models
A contrastive decoding method that contrasts a misleading prompt against a neutral rewrite reduces fawning hallucinations in LLMs, though most of the gain comes from the neutral prompt itself.
-
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
Relabeling hate speech as metaphor pairs (red/green, summer/winter) in prompts raises Llama2's F1 on a 500-item Bengali subsample to 95.89, though the gain is reported without matched test-set comparisons or error bars.
Discussion (0). Continue with ORCID to comment.