Pith. sign in

Paper Citation Record · LEDGER

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2507.12872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12872 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:38.745141Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:40:54.002702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:44.642321Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c8a5245-5b7a-4bec-a5f1-ef815bdc27a2 · outbound

This paper cites Towards evaluations-based safety cases for AI scheming.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Towards evaluations-based safety cases for AI scheming

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:33.388263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:33.388263Z digest=sha256:1711e168733e0f1c3171ddaa1b4bf3d343ba63670222e98cfdc6069546f0f017

Observation 3dd1535e-65ea-4ea5-b03b-12a6419c341d · outbound

This paper cites Sabotage Evaluations for Frontier Models.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Sabotage Evaluations for Frontier Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:33.510287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:33.510287Z digest=sha256:25ace04b73fcd8a9fd78df1587049b0d269655d55b57c156920cdc64c89ae553

Observation 5daba2b2-56c1-428a-9a14-c27588fa87d8 · outbound

This paper cites RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:33.918076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:33.918076Z digest=sha256:182f289bcf0ca71ecb5f83cb3c954b0ac6e044b4cd9241718a0ea932c596a86d

Observation a4c7f23e-f1b2-4122-a1ef-f42783b790ab · outbound

This paper cites Safety cases for frontier AI.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Safety cases for frontier AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:34.221645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:34.221645Z digest=sha256:42da2d34080961862461f5fd0cbdb9011a35a3b5932e098368c47754b66b058c

Observation 4fd43903-ab84-4a90-8136-6b47cfd3699f · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Discovering Latent Knowledge in Language Models Without Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:34.368752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:34.368752Z digest=sha256:8af7036d505b87541c20359dbe6fd5c244c1b9bc58d48c3813b83ecee0ead4e3

Observation eb126926-4866-4039-ba8a-20e9b42e8c9c · outbound

This paper cites doi: 10.1126/science.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework doi: 10.1126/science

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-06T16:39:34.536021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:34.536021Z digest=sha256:c85a8db0ea5038e8c9f17c193ee45777893e9a555deac5f96eadee26bb12803c

Observation 2f5945e5-cfdb-47ac-96b1-87d871d8d1b3 · outbound

This paper cites Safety case template for frontier AI: A cyber inability argument.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Safety case template for frontier AI: A cyber inability argument

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:34.797703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:34.797703Z digest=sha256:40c29d84ddda25698c73b553865ee78346ea39c04893dc4227ff297f85201af6

Observation 092ed59a-3920-4f54-a142-9417237a6709 · outbound

This paper cites Alignment faking in large language models.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Alignment faking in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:34.936205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:34.936205Z digest=sha256:aca59b4413ccfa80298dcdace6e827c2996b5f1c3652aa8b98987e9e9490b913

Observation d8251340-9940-4529-9b26-d5f2fa42defb · outbound

This paper cites Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.146041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:35.146041Z digest=sha256:25e4a164eb5516ffde906336315b8e1f18308aac97de4b2678d2ce64cd16886f

Observation f4514e2c-d7c9-4557-878f-44d85e69c1a3 · outbound

This paper cites An Overview of Catastrophic AI Risks.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework An Overview of Catastrophic AI Risks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.301994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:35.301994Z digest=sha256:b78d327f5b0b770f521f7f2513b74d3a2d6ca22ba18f7013824d9e2c53d076de

Observation 40ebb534-918c-4ad6-a927-557843c99c51 · outbound

This paper cites Facade: High-Precision Insider Threat Detection Using Deep Contextual Anomaly Detection.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Facade: High-Precision Insider Threat Detection Using Deep Contextual Anomaly Detection

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:39.585842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:39:35.608070Z digest=sha256:4d317f3eabda6aad0c2be5664e0480043b1b0b7128c45d2092c654234cb24c21

Observation 35a9ba68-1cf3-48a7-a135-5e6f71fe3865 · outbound

This paper cites A sketch of an AI control safety case.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework A sketch of an AI control safety case

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.743325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:35.743325Z digest=sha256:84c168edfb3c64400a4297ef88747b2ec34e7b7988a2f096c425e6f0c6e93276

Observation ca512bf0-6067-4e17-8342-e8b8d5e4dcdc · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Measuring AI Ability to Complete Long Software Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.884820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:35.884820Z digest=sha256:3148d6f7b07cde075f921b4d923964115e4ccb57b59209fc0747a99344992c6c

Observation 82d8d8d1-f84e-4a27-b3c1-ece7c72ba992 · outbound

This paper cites Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.071018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.071018Z digest=sha256:b90ac2f858afdc296f1f03a9458d0063ace60fc63b142c366aeb2fb60b64ca6e

Observation f295ef85-315e-4427-ad9d-383551dc10ce · outbound

This paper cites arXiv:2410.03768.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework arXiv:2410.03768

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.234981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.234981Z digest=sha256:b7f8a4d5ebe050f8c776a9bbccee8075ca3edf58fe25451c989c70c465371f5e

Observation ab9e34a3-8724-443e-82f5-200cdb825ffd · outbound

This paper cites DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:39:39.324600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:39:36.399493Z digest=sha256:f6fc4fe4f7560544dedf09c4a3f813d15de16ce55fb0bc31631b0b1b8fb93abe

Observation 4a7314a0-47fe-4c6f-85f8-3eaf24427fd8 · outbound

This paper cites AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.526571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.526571Z digest=sha256:91ec60c174ea9c0cff37d813cc6e19e592b3eac34e7c2a14aba133040a4d6474

Observation d196b84c-0b27-4bf2-a575-2e893d947be3 · outbound

This paper cites Large Language Models Often Know When They Are Being Evaluated.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Large Language Models Often Know When They Are Being Evaluated

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.667793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.667793Z digest=sha256:7a3da9e2d4f68150442f01be5b39080fad71a27b46eaeff7216d9e70fa5cfb8a

Observation 7579f224-aec7-4c7a-b527-bf7b68bc9725 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework The Alignment Problem from a Deep Learning Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.793291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.793291Z digest=sha256:9b4d1c95e5e7ada15b6267c3e6d9d8f5f71c403fc7f746c3b88b424b26467731

Observation 4e3323b6-488b-486c-a837-1b926db7d428 · outbound

This paper cites Linear Probe Penalties Reduce LLM Sycophancy.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Linear Probe Penalties Reduce LLM Sycophancy

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.959394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.959394Z digest=sha256:66b5a19cdd2d7605c990aab51cef0e7243fec795b0f6dfa43ae4e21cfa8a3a2a

Observation e107757b-5e07-42fd-a604-df71d28c0cf0 · outbound

This paper cites Evaluating Frontier Models for Dangerous Capabilities.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Evaluating Frontier Models for Dangerous Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.129859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.129859Z digest=sha256:e5e0856df33a86f2cf52ca38cec621ef01cee1526abe3f39b152b604b9225239

Observation 6f9fe1ac-4046-4eec-97e9-b50e3d3713ca · outbound

This paper cites On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.344120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.344120Z digest=sha256:07c409d6fe17869f6f007c10f8b318cd23ef7377823fcd4bf9fb2af7ed5ff105

Observation cbe3c3ce-fde7-47b6-9ef2-b51c5f46b7d1 · outbound

This paper cites Large Language Models can Strategically Deceive their Users when Put Under Pressure.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Large Language Models can Strategically Deceive their Users when Put Under Pressure

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.499516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.499516Z digest=sha256:4be83d283a14feac40ea8e2e39071a8c5e7cae63f85e72b70bd0b141b2c44986

Observation 254ec2be-34e0-49e3-abd5-1b93ff8b788c · outbound

This paper cites Melanie Sclar, Jane Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi, and Asli Celikyilmaz.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Melanie Sclar, Jane Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov, Yonatan Bisk, Yejin Choi, and Asli Celikyilmaz

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.594525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.594525Z digest=sha256:00332ad58640a858dac5c2d4182e3b119fc48bd413f43b39d10e3172e227ccac

Observation 8217a998-f671-41cf-85db-56b1dbc779ac · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.772633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.772633Z digest=sha256:4732e7dfab649345f66fd8ab658992464494cdbdcd417ed365c9f2469b25a82a

Observation 06a423fb-b5be-4a4d-91f9-b1f054111d87 · outbound

This paper cites Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani, Fedor Ryzhenkov, Jacob Haimes, Felix Hofstätter, and Teun van der Weij.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani, Fedor Ryzhenkov, Jacob Haimes, Felix Hofstätter, and Teun van der Weij

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.919794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.919794Z digest=sha256:4d7b5ecc8221ccff1a6cd8c2196aa4e73b79c772790ab1a9688a8616c242cfd5

Observation f6142e7d-acb8-4907-bdbb-ae1fba6c8cdd · outbound

This paper cites Oriol Vinyals, Igor Babuschkin, Wojciech M.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Oriol Vinyals, Igor Babuschkin, Wojciech M

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:38.000260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:38.000260Z digest=sha256:1f4eafc14dc6c8b943dc19d1627966b99c2fe5aa1de829e0de453d99da2a86eb

Observation 0c2aa3db-74cd-4152-8fa1-dfcef47d83d0 · outbound

This paper cites On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:38.443628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:38.443628Z digest=sha256:6b71ac6f51c57ac92e053b70533c12caad69e7f51696c8e5c815012ebd760ebf

Observation 88aa5c2d-366c-477e-8608-eb2a8a0693c1 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:38.613609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:38.613609Z digest=sha256:658201cd5294cf89bef39fa3dae386dea8452023929f186b8d1317dad040e25b

Observation d92ccab7-940a-463d-a334-18a3d93c4c69 · outbound

This paper cites trusted inquiry.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework trusted inquiry

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:39:39.865103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:39:38.745141Z digest=sha256:7c7b4dee4e650128bbdda6ba1519e20683ab91d09222d9e32d1b378d8d678b25

Observation a6467fe5-0cce-4521-b018-c2720d3a5487 · outbound

This paper cites doi: 10.1017/S0140525X00076512.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework doi: 10.1017/S0140525X00076512

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.217624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.217624Z digest=sha256:dc453f2209521c2e3cc7921dafe9ec05c336f12b7064e3a6222435db24194ab2

Observation 54ff7abf-df0b-4ce8-84da-b5785de4165c · outbound

This paper cites doi: 10.1038/s41586-019-1724-z.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework doi: 10.1038/s41586-019-1724-z

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:38.187897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:38.187897Z digest=sha256:e57d908d7d94cdd9939ec7f90a575dd5b0e476b46e6b240766fe1c144b6038f0

Observation 45ba4898-9e81-40d1-8001-4cef1191468b · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.476815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:35.476815Z digest=sha256:b97a3f02e162638ff0bc82bdf6cec506c3016d74e2f4b8d84dcc6d2f66b2b90b

Observation 9dbc06b4-a131-4486-9e6b-75de0ea0bbee · outbound

This paper cites Emergent Abilities of Large Language Models.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Emergent Abilities of Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:38.354084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:38.354084Z digest=sha256:2b4f0f7e5fe9ae831b04b88f0fb5388aac045837ef44f03f907d57829b684886

Observation f5b9cf08-8480-49e9-9190-96519907496a · outbound

This paper cites Taken out of context: On measuring situational awareness in LLMs.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Taken out of context: On measuring situational awareness in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:33.641786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:33.641786Z digest=sha256:dc65df28c008604fe5006e4eca2fb9e3cc9fb3d4b19d91024263883d0df8b781

Observation ed0ebd9f-1d21-47c4-b787-933009b79d1d · outbound

This paper cites Anthropic.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Anthropic

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:39:40.048618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T16:39:33.247242Z digest=sha256:eaff598669517885e8d3fff9d49c116eec0e6c719a72f202c7d36c11da3185ba

Observation 86c9ba4b-f572-46cf-8cae-43602dff42c7 · outbound

This paper cites Ctrl-Z: Controlling AI Agents via Resampling.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Ctrl-Z: Controlling AI Agents via Resampling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:33.772218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:33.772218Z digest=sha256:cd0ad6fde7b724a33fe63ec6a4e798f6d18dd8637f0c0a4cd3c8c12597ddf8c2

Pith citing papers

Observation 5afbc83d-aa33-4eae-bf22-40991007b544 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.643863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:e4b911dd80dc5dc5a2759d3f7551ee09a9de0d46aec5ee99eb74dc192e9be53a