Pith. sign in

Paper Citation Record · LEDGER

On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2411.02306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.02306 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:19:06.259747Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2413bf6e-d7fd-4860-8a16-874b946cae07 · inbound

Observation Interference in Partially Observable Assistance Games cites this paper.

Observation Interference in Partially Observable Assistance Games On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:19:06.259747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:19:06.259747Z digest=sha256:3777527eddb1cf518a85f4aa23c2f0de5d1eb93fe595883033d68b43ddacf1ff

Observation 677c6706-98f8-42a8-819e-52000ef3335c · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.740291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.740291Z digest=sha256:37b894efed5861a5432648ae821ca2096b78efc4a6b61a97f6986b48ed98291f

Observation 5cb3d740-ddb5-48a6-af24-12f6d1506173 · inbound

Why human-AI relationships need socioaffective alignment cites this paper.

Why human-AI relationships need socioaffective alignment On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-09T11:53:43.873077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:53:43.873077Z digest=sha256:b9f2db3e3343f38ffcf3c2eda4fe886005b5005d395a869903dd785c7d3eb709

Observation f988d7e4-0356-480d-96c4-767364624d4e · inbound

The Lock-in Hypothesis: Stagnation by Algorithm cites this paper.

The Lock-in Hypothesis: Stagnation by Algorithm On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:29.471740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:07:29.471740Z digest=sha256:e41d529f110c5ab1c65a69d77093088960458561c924229ef0390fb17151ba10

Observation 0c2aa3db-74cd-4152-8fa1-dfcef47d83d0 · inbound

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework cites this paper.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:38.443628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:38.443628Z digest=sha256:6b71ac6f51c57ac92e053b70533c12caad69e7f51696c8e5c815012ebd760ebf

Observation 1a9a12ca-456d-490e-a71b-77b9dc424b9f · inbound

Mitigating LLM biases toward spurious social contexts using direct preference optimization cites this paper.

Mitigating LLM biases toward spurious social contexts using direct preference optimization On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:33:14.711288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:33:04.433907Z digest=sha256:17b50b85ab15e60e4a86ec40db6d7e5a8b15a16f2c57fa505c3d133d002911f6

Observation e7853648-c97f-4152-b45e-93f4915b0e0a · inbound

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants cites this paper.

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:47.497987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:28:13.317630Z digest=sha256:b71ec8e2db1362d7e0ba6497439cc5be4ec5b7b5415728e29397b849d68c177e

Observation 6f416061-052a-4a36-ad3b-73db504dbbe3 · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:26:26.433528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T04:15:24.919355Z digest=sha256:5290fc9aafc376c0513380c03f1a7a0a0292d061e515564831997d5fa1574c86

Observation 34fee9e2-f2cf-4589-9372-4ee2e14d56ed · inbound

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning cites this paper.

Structure from Strategic Interaction & Uncertainty: Risk Sensitive Games for Robust Preference Learning On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:08:05.074380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T22:03:05.102274Z digest=sha256:93b4e3a8ef3914befde34984c0fe83f8eae56bbbf90c0742b7a30b49ac4e802d

Observation 3895b28b-088f-46cf-9173-a7d42d07ca6f · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:22.314555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:94f325eb6e39f763f2b04cca16885c7683e545a1f1e196fbeef1735b3b14b216

Observation e0525972-eb1f-4cbe-b8cf-92f9b5659cf7 · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.438931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:89a2fcb039caf6ba28c4275da499abb9f63ca5bbd2fd8a9839d269ac613a490e

Observation 32ce9421-a677-49fb-bb53-5592957b5e3a · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.065439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:230ab45fb781000b98436e233ef65bc3e40bc655f393fc50fa324c52e6cde3b8

Observation 5af3ff9c-5371-481e-aade-f6878d2db672 · inbound

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing cites this paper.

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.587162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T06:17:01.173495Z digest=sha256:b102f744a9b9859feb227b1bf865d7698227dbb60c74bfc49f57bf89ede2685d

Observation 05b8766f-9d6d-4619-8d5b-3c0d94259f68 · inbound

Against Proxy Optimization cites this paper.

Against Proxy Optimization On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.629443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T08:26:50.595370Z digest=sha256:74e3f1e43298043ae2592db03c9c298aba70374a0dee9f2c113b1b6e0c8ccb62

Observation b8249767-90a5-44c7-bcc2-52b47f825175 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.649060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:84b971838f8d916bdabe5b83be4bfe59a65fb22ffb990329f6d883fcf2a6cfa2

Observation 948a15f7-17d7-4f76-8b0f-5b311c1ea7e5 · inbound

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being cites this paper.

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-31T02:25:21.866266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:25:21.866266Z digest=sha256:afe754854dade6c0adbb8425fe5fdd0b6b584f2809d3dc59daeb2de3a3895eac