Pith. sign in

Paper Citation Record · LEDGER

Compromising Honesty and Harmlessness in Language Models via Deception Attacks

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 5 inbound Pith citation observations for arXiv:2502.08301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08301 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:42:43.602065Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:09:27.931667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:07:38.858520Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30a7df6c-b566-479a-8c92-ce6dcaa8b469 · outbound

This paper cites deception attacks,.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks deception attacks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.684083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.221711Z digest=sha256:7f1bb7962f009ad1287105ecc31c43d45dd8bcc330039d4040676d5fae94eb8f

Observation f53513aa-31b3-4962-8c6c-b5e4dac5edd3 · outbound

This paper cites AI Safety in Generative AI Large Language Models: A Survey.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks AI Safety in Generative AI Large Language Models: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.350763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.350763Z digest=sha256:2019ee14b47a9ca235226d846da1d255a3a916df70f46adf9bf1dc0c0ca39581

Observation b8e16009-3e6d-4d5d-bb35-153f11c90f23 · outbound

This paper cites (a) GPT-4o, (b) GPT-4o mini, (c) Gemini 1.5 Pro, (d) Gemini 1.5 Flash.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks (a) GPT-4o, (b) GPT-4o mini, (c) Gemini 1.5 Pro, (d) Gemini 1.5 Flash

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.699526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.213880Z digest=sha256:28d59c49422e009ea8de7920a6fc2e3e2214cf243b958e214906e6447a5be4b1

Observation abee7819-170c-46ff-ab91-a9799ca0f8ab · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks AI Alignment: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.343857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.343857Z digest=sha256:67ffbce650563ed7730ac5429d90d02eed338b5c853a6c199dcb16581e72f2d6

Observation 2426deb3-3272-4631-bb27-b051e40c7d9b · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-Tuning Language Models from Human Preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.356519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.356519Z digest=sha256:06b380e6ba69989ea1d921449641bf5e7ba4d312500e494f95185f29d55dbb7d

Observation a72a9211-28f1-47c6-a8fa-832aa034fc41 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.362622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.362622Z digest=sha256:b341926cb06d520d4dba4e88ceadfead41f033ad4a6527b1f01dee6bb983fce5

Observation b03ff4b1-7490-4fce-a9b0-65d5f4f1bf7a · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.381394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.381394Z digest=sha256:b8595cbb6d90dca371de6b5125263101a5aedac0e741aa8d9e099fb41829155d

Observation 23fc67e0-3355-4317-b1f5-14535894b972 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.386584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.386584Z digest=sha256:2ff6930871a34db0655a41b2bf144b81812bb0624f7d07871935f56c1886dd6d

Observation 33be2679-36b9-4821-ad97-92404033e70c · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Jailbroken: How Does LLM Safety Training Fail?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.392787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.392787Z digest=sha256:8bb5178577d3e5a1592a817c8a0d3a3f03e2377c2f660febce23f206cd35d105

Observation ce9e6ed1-a7b4-4f5d-ba97-53a5f8b8b45e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.400191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.400191Z digest=sha256:b5f5ac8f4ec9e561ab06ed6f288f26c4d60984be8accd088a2cb50657413bf6c

Observation d724c645-9cf2-4ecc-9449-c06a54b05d5c · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.405483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.405483Z digest=sha256:c358d981ba1bdcc2b80338a436dad8ba4eb2908854a2bcc72b3414ca71dcc4dc

Observation 69cf55cf-428e-4b79-8d6d-91b5bd5ca8d8 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.410849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.410849Z digest=sha256:0d7b374a227e53a001deae9f7aab4a591c0f4f9d97091f19d66b229926f855d3

Observation 12750872-ac46-4bd2-bb57-f91686ee783e · outbound

This paper cites Mapping the Ethics of Generative AI: A Comprehensive Scoping Review.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Mapping the Ethics of Generative AI: A Comprehensive Scoping Review

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.667049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.416213Z digest=sha256:d6dfe3f0dbad93b44773a78e9d9430e93f5d8f3a219f9e547f045abb9814f1f2

Observation e0b1aa67-3e4e-4af0-b51d-799c32909168 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Alignment Problem from a Deep Learning Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.420807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.420807Z digest=sha256:f45d47cbe58596d74bbe4fdc7b119ee0c8a59aef19f1b2613fd40f155b10a677

Observation 55e42007-e175-49a0-a43f-2f0d70b6c994 · outbound

This paper cites S., Goldstein, S., O’Gara, A., Chen, M.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks S., Goldstein, S., O’Gara, A., Chen, M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.649151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.425803Z digest=sha256:03a662198b01df815e172aaafebb11dc0a568e16b8fc440cae4a1eb760a23c93

Observation 98e5e8ab-fb70-458b-965b-0859d744ac80 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.431086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.431086Z digest=sha256:3419d912db2ef6a4613dcfde6fe1c56cb534e5dadce32a48c68fee68a981ea13

Observation b307d4f7-9636-4544-bb80-90daabd2a075 · outbound

This paper cites Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.435854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.435854Z digest=sha256:6232e4f729ab05745ef670e492c8583253ad4870910e2a8a2cb28d7be52e2810

Observation 12bd7929-6ffb-441e-90c7-32071a6903a5 · outbound

This paper cites Scheming AIs: Will AIs fake alignment during training in order to get power?.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Scheming AIs: Will AIs fake alignment during training in order to get power?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.440989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.440989Z digest=sha256:876e370951b78edc177e2df6c4a1532660243bd3bc65c93031d117a2dbb2d738

Observation 7e83aa47-f615-4f15-9320-163967861c45 · outbound

This paper cites X-Risk Analysis for AI Research.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks X-Risk Analysis for AI Research

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.445722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.445722Z digest=sha256:548585bae71f3f023671f73dfca60d087f13e6b4df061b2228e0299dc59f4e6f

Observation 7e1b6a93-7996-4fbb-af4b-632e13c3489a · outbound

This paper cites Deception Abilities Emerged in Large Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deception Abilities Emerged in Large Language Models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.630371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.450559Z digest=sha256:4da9170c77dcbb64f32a90fd80082d13cd83033c50ef97ec70bb464b14d59f3f

Observation ce3383a2-b24b-4810-8809-72e57343aa76 · outbound

This paper cites Alignment faking in large language models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Alignment faking in large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.455240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.455240Z digest=sha256:588ac9460646ebd3d9285e07017832635f892eb2926b959ae21243e4e8ae798b

Observation 8274109e-2066-46e9-9ecb-19047f44c99c · outbound

This paper cites Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.460514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.460514Z digest=sha256:364527b478aa78fe4ab9e2992afaa66ef0f2a8a73af1635e4b8e3c53868f1326

Observation e04a58fa-88ce-480a-abc3-efc2237f97a4 · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:42:44.612703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.465366Z digest=sha256:eb0dabf1a95ae964b941007520b18bc333c1d2228be6a22408b47b5d77f370d2

Observation 05f058c2-49ad-45d3-a520-81295286a87f · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.470194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.470194Z digest=sha256:db2124988b6dd1112b664c2a7a2615e2b6cfa97d9f13a4d75df81c4de7309667

Observation 8ba51696-cc2b-45bd-98ac-6316dedf5d0e · outbound

This paper cites Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.475213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.475213Z digest=sha256:3339c14ffd06c9dbcb015a83639717884e467112faf33fc090dbfcbff774b443

Observation 71e8a86e-b8be-40e0-a00b-60052ead95d4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.480254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.480254Z digest=sha256:0425beccae3fdc849a76c1cd766d21419ec165aa9bc62c7700c20d11b5b8f277

Observation 040b0cef-6164-43aa-bc92-5ac51ea9a613 · outbound

This paper cites The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.485261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.485261Z digest=sha256:40c04c0e871f2f9506691052537e498de61c1a6a6b28fab993df64ed8cb056a6

Observation 1feafa60-095e-471d-bc14-54ee3af548ab · outbound

This paper cites GPT-4o System Card.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks GPT-4o System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.490299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.490299Z digest=sha256:ad794d39d052fdcba778a2a3e99644f1bbcc80caba02cda69c750c922270fb86

Observation bc8bd510-58b2-4c83-8b13-2ebba885cc58 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.495381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.495381Z digest=sha256:a50b94d7c298197ccdca33df5899c34875a471d221ece450d0fe2a06dc075b6e

Observation fefc191f-321b-4aed-aa94-e47869e802a3 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.500457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.500457Z digest=sha256:391b5022c7fafe1c18ff81d3874a0b3fcb05e7cbd9bb659a572f942250832b75

Observation d1c608ac-143c-4e23-8001-10d4687a0371 · outbound

This paper cites Mitigating the Alignment Tax of RLHF.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Mitigating the Alignment Tax of RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.505543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.505543Z digest=sha256:592e98dfce599a2f26ce2749683fcf4c50e8cdf94a0123bda995e0495a4a9f90

Observation 0bc28185-bb8e-4a92-aed4-cb1fe526fdfb · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.510437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.510437Z digest=sha256:c88599792b0dd9f9e3bc97248fb4793d9b0119d1bda2b5e2dabc7e63b82598dc

Observation 37e0162d-52a1-4d5e-a91a-e77dcc0125bf · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.515096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.515096Z digest=sha256:603ca10b578afd25fc1341fe72f0f69d20f3e1618b7996c239544d0eae300a39

Observation fae777b6-1fbf-4673-bc0f-c4676c159bbe · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.520707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.520707Z digest=sha256:d246c4ebf2657aaf6f31d59dfd6a2761d685f6b8caa39d19793abad936ce7391

Observation 69100820-9d57-4bab-ae0e-9982ee853213 · outbound

This paper cites Large Language Models as Misleading Assistants in Conversation.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Large Language Models as Misleading Assistants in Conversation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.525731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.525731Z digest=sha256:b409a5ddc55825a69d1e48b837e7cf50db876c247b5944b438a8efe734a7b04c

Observation e5891618-7dd3-4bf0-b0fb-231de967a490 · outbound

This paper cites OpenAI o1 System Card.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks OpenAI o1 System Card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.530889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.530889Z digest=sha256:467d26824e101c2ea1f43b6bf10710db829fb721a36a8b129a328e16e32b2ec1

Observation f4c3cf1d-52b0-4968-83b0-138bc8c69b4e · outbound

This paper cites The Llama 3 Herd of Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Llama 3 Herd of Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.536036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.536036Z digest=sha256:7ace1ada0628f8ffefe2db5faf844e327cf11ebdd4bc905521754f77e0caa030

Observation b7f88c17-f783-49cc-b2d8-562a0a7e9952 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.541142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.541142Z digest=sha256:f53dc334bb69e9022e585cf7aac6bd12bdb69db2a6c310b83a41213a64057464

Observation d4d64535-26d5-4d38-9ae9-8e6cccb923ed · outbound

This paper cites Claude 3 model card.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Claude 3 model card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.596621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.546966Z digest=sha256:c35f39c41f7ac4d5cbb54ee21c08c7995af20d80ba6479853ecd7918a55fb08e

Observation c803a1f3-b23a-424c-9c33-cca2ba79e142 · outbound

This paper cites Do Large Language Models Latently Perform Multi-Hop Reasoning?.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do Large Language Models Latently Perform Multi-Hop Reasoning?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.551459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.551459Z digest=sha256:2bc4ddad3523dcdb321c508d91bbdfd0ee06ec1567fb10f89bdc0cfedeac735b

Observation c87d9c3e-ba6d-4aff-aabc-d4b7d2daa036 · outbound

This paper cites Human-level play in the game of Diplomacy by combining language models with strategic reasoning.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Human-level play in the game of Diplomacy by combining language models with strategic reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.581368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.556376Z digest=sha256:eea8c6db6fd545672b1664021ab5ffb3157bda245ebfdf4f34629f351aa50ff0

Observation 1ebd489e-c5c4-468f-95ed-02ced5ed45bd · outbound

This paper cites An Assessment of Model-On-Model Deception.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks An Assessment of Model-On-Model Deception

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:42:43.762712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.561395Z digest=sha256:a1885918e2932ee1a5e1db0120383c5f719b51df837c097d26c3f24178292cad

Observation 663d5dc7-ed42-4e2f-876e-021342c147b7 · outbound

This paper cites Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.566302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.566302Z digest=sha256:774252eb43aac8255f29e891f5a053d7cf00609790caa7bafaa95445afee6831

Observation b9848e39-91b4-4261-8ac9-02557fdf9b21 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.571564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.571564Z digest=sha256:5a0d84c1c55261f6b771e0912cb9d8613610fc202b89bfcdb8cf5814b1d74431

Observation 5b9394cf-220d-493e-9878-6061ce5c43ad · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.578787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.578787Z digest=sha256:9e7cc5670fdb30848c9cfbd7ff2f3b8f39ca09212eb50bc91a5b6a46be0a9896

Observation 1cc15828-65e6-45bb-9b97-2908bc6a33fc · outbound

This paper cites Tell me about yourself: LLMs are aware of their learned behaviors.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Tell me about yourself: LLMs are aware of their learned behaviors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.585155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.585155Z digest=sha256:11676015b7c315c0e6d0a2d50fc9d0b9ec7000a7a93f864c2e50fad693501b4a

Observation 3d6f8215-9986-4c9a-b101-f4fbde9565f9 · outbound

This paper cites A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:42:43.655382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.591256Z digest=sha256:a8b4e97f0a545bb7c6689f9ef066142d74109f603ee94df38e3d78ac4dd807c9

Observation b0f1d33d-59c3-4970-b818-23f5f4981dfb · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:42:44.564606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.596383Z digest=sha256:ecc835643f1a48077080b80fc1a4d5527b0743e4e381e72b6ab5e7cbccec9a97

Observation 09fba7d2-a817-44e2-a1e3-4cd9dddc4705 · outbound

This paper cites Italy” , “Queen Elizabeth II.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Italy” , “Queen Elizabeth II

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.546282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:42:43.602065Z digest=sha256:8d0717f69ee59c4481a41baa718212c5a661c1063dc54ee091d144edd40979d3

Pith citing papers

Observation c84cbc0b-eb40-4d79-a65d-248609831681 · inbound

Model Organisms for Emergent Misalignment cites this paper.

Model Organisms for Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:28.223256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:28.223256Z digest=sha256:4cc3ee8c5188f7f760e95fced5c73339ea4c9eb04a8dd681650b37606671f66d

Observation 4546a6a0-f0bc-4891-8038-61c728b71d2a · inbound

Convergent Linear Representations of Emergent Misalignment cites this paper.

Convergent Linear Representations of Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:27.931667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:27.931667Z digest=sha256:418bf76ea3b3b9c119a32b5ee2a6fb46512a0e7291679d7a454c051ef8e7774a

Observation 68388b80-fde5-4ba3-8acb-5316fda5363f · inbound

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment cites this paper.

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:42.234545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:42.234545Z digest=sha256:449bc82ec42f5536378f7c918df76aee87a0277b9a58b0dfa5c0cecd78a1aa88

Observation d5ae6605-c1da-416d-8be9-7433cd24a848 · inbound

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating cites this paper.

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.004832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:03:33.199645Z digest=sha256:de4fac31c128f57228110ffb12cd9d8309b4a97bbfb1ab33dc83374ac5b1ed24

Observation c51d0efd-9cd9-4a63-b8ae-1f20f9a247f3 · inbound

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment cites this paper.

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.860058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:26:14.457195Z digest=sha256:75189ff2767acfc3c74a072a9a4a497d4b8ab14b69b6c8697cf03bd23ee05d04