Pith. sign in

Paper Citation Record · LEDGER

Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2311.08596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08596 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:43:27.307708Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc3d6c4b-c7a7-4956-9545-80b57ffb5c06 · inbound

Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies cites this paper.

Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T04:43:27.307708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:43:27.307708Z digest=sha256:b97e26051dc6ba803d0e9d40eb3172df1276e69dcf2e64c632fd501bda156b42

Observation f1d13927-1dc7-4c22-9040-94661f54fb49 · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.232010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:c3b6ca2901d90c20df1f6f8037a3ca95fc43b62d1ccdab31fd5b3c84e2f42a04

Observation bc8217e7-49bb-4a50-85c4-b5747c275eff · inbound

B-score: Detecting biases in large language models using response history cites this paper.

B-score: Detecting biases in large language models using response history Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:49.174083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:33:49.174083Z digest=sha256:78973b7256c946f90a14da2fb2221bd05e55372b4e4594c31c829070b393463d

Observation c4df5e8a-0f8c-4991-89c9-85255de74185 · inbound

Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems cites this paper.

Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:25.333973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:25.333973Z digest=sha256:75ffc7000a4b2f6a3b29682df91a86cb794f1dc03239d2811480adfd0e7ff1d7

Observation 1e87feb2-7423-4e88-8b1b-29f2c3fdc83d · inbound

BASIL: Bayesian Assessment of Sycophancy in LLMs cites this paper.

BASIL: Bayesian Assessment of Sycophancy in LLMs Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:56:52.000793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T21:55:45.714195Z digest=sha256:c0953a7c0202d2f47d6a4fdd91e1c181e86a2f63c16080d17675f078fdd8b4e3

Observation 5c16d096-67af-4732-89ab-b2a341a18c89 · inbound

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI cites this paper.

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:07:58.764297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T14:03:49.595868Z digest=sha256:5a37b0cb861a2ce140d09932e3bc5a77cb818f0ffde79ce3a3be23b3ddf16ab1

Observation 5558a29a-fe74-4a63-9907-b2d985e4f019 · inbound

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care cites this paper.

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T08:39:30.412601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:39:30.412601Z digest=sha256:1552193e3ae33545a6c9c71eabaf315ca9e72b3ce6bc9369016221ad396ae1fb

Observation c482b472-01bd-4493-ab47-f2ca510ea8c5 · inbound

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems cites this paper.

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:13:13.995294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:09:14.217158Z digest=sha256:83b6104a7997f8b5604452a63daaa1c83c084e4206564f19e393c17ae819347b

Observation 18e89a7e-ef12-4a18-9363-609cf5f14d5c · inbound

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems cites this paper.

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T16:57:39.838460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:57:39.838460Z digest=sha256:39fd7d13309609a3731c05e95680eb7ebd895873898b66a6e2fc83df0b5b376d

Observation 0ae3a3d9-84c2-4bd2-a46d-96905dd72ccb · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.414046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:b06c2a904819b56774f18fbc42f893a012d58855263691e179b08d57766c9143

Observation ef697d0d-7b1c-4878-80d9-aa4b018aad85 · inbound

Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts cites this paper.

Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:12.899263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T10:17:35.694775Z digest=sha256:0a2095c6ee94c7986eee4ec9263a107db7f931101b217ca3a2972055c486f031

Observation 535b8f9a-eaf0-4bfe-971f-a1c5c592ab12 · inbound

How LLMs Are Persuaded: A Few Attention Heads, Rerouted cites this paper.

How LLMs Are Persuaded: A Few Attention Heads, Rerouted Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:26:24.301302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:18:09.600353Z digest=sha256:b269ca835c090a4477718655f239795e23c1d12cc820e09b220c33b63eb5b79e

Observation 1c1c0c47-da5b-4c5b-a3e6-a4f793afc945 · inbound

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct cites this paper.

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:51:18.405646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T08:47:51.191086Z digest=sha256:91eeb87ccfcd51ddddd388ae23fb8d98c385d6e0b35e75c049c7725261972a01

Observation 902fba5c-3174-42de-b380-d99934bb858c · inbound

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience cites this paper.

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.449174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:29:12.725865Z digest=sha256:77dfeabad1574cc8de380bc8b56226f084c0c79392cbdd0b5e94dfc1a75238f4

Observation 4179136a-eaea-4438-9913-19f8746a1566 · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:07.569477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:3c003e457912e6856c1911d262e3f16e06ef0bd2b72f9868c8c1c7182a0dd99f

Observation 2a291b7c-4e4b-4143-85b8-bc784d7045f3 · inbound

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks cites this paper.

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T06:02:15.946864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T06:02:15.946864Z digest=sha256:f313465eb7c8141a5a1acd893dc3b76a18b5803cfdc1a5bd22d0461e5a928f54

Observation 11a93628-9b4a-4d3f-ad51-483a63abe15b · inbound

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs cites this paper.

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T05:39:51.084130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:39:51.084130Z digest=sha256:c07c3a77aa141ad94d3feefa9f4a4b5028268f657d2b8a0dacebbffd4b154a34