Pith. sign in

Paper Citation Record · LEDGER

Mitigating Deceptive Alignment via Self-Monitoring

As of 17 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.18807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18807 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:10.795238Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:14:22.854339Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 83677165-82a4-4664-bd97-21887015e19b · outbound

This paper cites Introducing openai o1-preview.

Mitigating Deceptive Alignment via Self-Monitoring Introducing openai o1-preview

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.690887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:03.992286Z digest=sha256:5ff3df77b650b9cb68f68890c07cc5d35f365097cdb1acd00ac0fa07033f4155

Observation bc6cad5f-5a5d-4302-82d2-65ce6a296a0b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mitigating Deceptive Alignment via Self-Monitoring DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.060422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.060422Z digest=sha256:0052ae9f64a1064ebb470cfd35d6582609f58013322ddbce11c14011dfc18afc

Observation 51cf49de-8586-462c-bddf-9c41241e9cc2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Chain-of-thought prompting elicits reasoning in large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.114784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.114784Z digest=sha256:ff9cbc45d47f7a99fc0170ce6c5b8abf0a37895b18327639177ba2268a22e939

Observation 09c55401-661d-49d0-848f-bfa1201258ca · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Mitigating Deceptive Alignment via Self-Monitoring AI Alignment: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.196418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.196418Z digest=sha256:cbc981889a7b1f0e5f89a6767784784b50297d06f4a1d840bc5e07660407ee62

Observation dc773ee9-4491-4355-afb5-143fe9e19624 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Mitigating Deceptive Alignment via Self-Monitoring Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.260470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.260470Z digest=sha256:5ca29318d4387eff12ed139014c286be25bd46b32eed463b33f39e5acf2090c1

Observation 80a65e4f-f2a7-4382-b739-6637b0c24c72 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Mitigating Deceptive Alignment via Self-Monitoring Frontier Models are Capable of In-context Scheming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.319823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.319823Z digest=sha256:1990031e2ea04df59310192044754e98ea4469d2a1e0679bf9db510b329aeee4

Observation 0c2bc9d3-9b3a-4bd8-a2df-7f37c0a94918 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.428012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.428012Z digest=sha256:beecdda40bd728fb3c33bba6a722ee8a10e147f13393d5485eb651376935a50f

Observation 904a83a5-b850-4e07-b3f5-645ad7b05482 · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

Mitigating Deceptive Alignment via Self-Monitoring Language Models Learn to Mislead Humans via RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.463941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.463941Z digest=sha256:b86b16bf038e7918e26f23b9b018fcc3be033f8832bf1aa519b5fcde3dfb9488

Observation 3bda92d4-c08e-4a48-90b0-53f4e7113efb · outbound

This paper cites Managing extreme ai risks amid rapid progress.

Mitigating Deceptive Alignment via Self-Monitoring Managing extreme ai risks amid rapid progress

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.536207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.536207Z digest=sha256:6a4029d45f6693362e7e8e35f9ed080aced1503351dd0686b664f56dfb5dc4ef

Observation d1c7d851-abb7-46c1-ad89-0635acd982e8 · outbound

This paper cites Privacy risks of general-purpose language models.

Mitigating Deceptive Alignment via Self-Monitoring Privacy risks of general-purpose language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.514680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:04.735429Z digest=sha256:a8e52314f9f61bc4bb3de05ae55061390ffa3d7004a262238f310041d24c1c8c

Observation 098f5cbd-18e0-4689-9dbb-e7359454e794 · outbound

This paper cites Frontier AI systems have surpassed the self-replicating red line.

Mitigating Deceptive Alignment via Self-Monitoring Frontier AI systems have surpassed the self-replicating red line

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.787401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.787401Z digest=sha256:8dd77940365953006c2e763cc511b0b7d4152489d2a7851dddf863a30c8a36e4

Observation f0b4bd89-2800-4f79-8679-5371b3c50fe4 · outbound

This paper cites Alignment faking in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Alignment faking in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.868190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.868190Z digest=sha256:ca44146294daf49a6c4ac5ba459316e0b9347a4d3bf3aee290d8d4c154c03d1e

Observation 2fa7d15b-f588-41bb-b358-f85bac93806f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Deceptive Alignment via Self-Monitoring Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.987930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.987930Z digest=sha256:7b9693335da4e93626bb123dc10a7b153353ea393f4d3c6945b00f65b65e6bf9

Observation 26c081fb-2fb4-4bc7-a133-8512d908309b · outbound

This paper cites Darkbench: Benchmarking dark patterns in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Darkbench: Benchmarking dark patterns in large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.182736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:05.047058Z digest=sha256:63fc075164d2891d203a27deaa72713756ccc82f238e41168d77b1f5c643c2a9

Observation 40edb399-1f95-4346-86e6-090ebf81128c · outbound

This paper cites Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?.

Mitigating Deceptive Alignment via Self-Monitoring Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.120166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.120166Z digest=sha256:88a5dad61607acc018d46a75267764b7c3eddd8292df0ef8ac582d2d4431ce09

Observation 673ccd70-0688-48b4-b0f3-f48a23bd1dec · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Mitigating Deceptive Alignment via Self-Monitoring Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.188094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.188094Z digest=sha256:60d4b5f389070581e06f6dd31434c35d847568c9987020ecce2c30b3446599c5

Observation 8d4e3e7a-8f66-4be0-b3fd-1477604c2cd2 · outbound

This paper cites International AI Safety Report.

Mitigating Deceptive Alignment via Self-Monitoring International AI Safety Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.234059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.234059Z digest=sha256:27975bf6bb364ae98793baa6c6af009dc4f0ed8222520f4909dae2ad98909c42

Observation 8ee0c132-066c-4f33-832b-40e96438b21d · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.

Mitigating Deceptive Alignment via Self-Monitoring Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.260854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.260854Z digest=sha256:60d8e4da983c5e332684f40b35445183b311b6c4f982add83426c7903749bdba

Observation d9afa334-7493-405b-b63e-a82a78f5d73c · outbound

This paper cites Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.298607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.298607Z digest=sha256:707a5019d4a289ab568fb7f45b11d779f63cde800847ea9451e7317711925aae

Observation b0e2b95f-069b-4d5f-b70c-6d85580017cf · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Mitigating Deceptive Alignment via Self-Monitoring Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.384371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.384371Z digest=sha256:d44357e4a318d53153f2d9ed4afc9a536bfeb5e7481bfa361357d6fb61b69b1a

Observation bea15da7-e936-4cbd-84b6-7fe78ea1b9a9 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.477915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.477915Z digest=sha256:658ec2435cd1c4d4510284690e038c999ff62e76f4db7e04abb4874fb780cb7a

Observation e12a21ad-c180-41f2-a85f-4a0d1874be96 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Mitigating Deceptive Alignment via Self-Monitoring Markov decision processes: discrete stochastic dynamic programming

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.595053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.595053Z digest=sha256:ddc47e9a37a4f9197d681be43b1f8c8f7402a393d4304f2e228ad35d7a135282

Observation 8694ca25-0e87-4150-a6e4-bd758a5fc61b · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Mitigating Deceptive Alignment via Self-Monitoring Reinforcement learning: An introduction, volume 1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.716513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.716513Z digest=sha256:7015aac2163628ed2728e3fc9d22abdf6c1205dbc40736f096878d3d93bea54c

Observation 02639ee0-cd8c-40c4-8bb0-31684ae7ceef · outbound

This paper cites Training language models to follow instructions with human feedback.

Mitigating Deceptive Alignment via Self-Monitoring Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.874104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.874104Z digest=sha256:4775b393577aa055f750b40ad3bc2ccdb8dd3a7be1f93d22244dcfd6af8d88dc

Observation 7e9cdb46-9f25-447e-ad2f-6c6f25721ab5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Mitigating Deceptive Alignment via Self-Monitoring Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.978324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.978324Z digest=sha256:768782d32960b2d9699bb1a7ab01971da7b9a0f1f67906af5917dcc7c9ace848

Observation 1cb8ea0a-c44f-461f-a6cc-02ebeb23f7bc · outbound

This paper cites Handbook of constraint programming.

Mitigating Deceptive Alignment via Self-Monitoring Handbook of constraint programming

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.898223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:06.196688Z digest=sha256:58f68975ba02ae052388352415a891c3cc083da6e75b5ccbfdc18f72897aaeeb

Observation c4c02e54-93f3-4779-a680-44c8d90d28b3 · outbound

This paper cites Defining and characterizing reward gaming.

Mitigating Deceptive Alignment via Self-Monitoring Defining and characterizing reward gaming

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.349381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.349381Z digest=sha256:6f590f18daf221c0b6978b29d5b02ba033202d357fb835a47bae36e342fa264d

Observation bbc388a2-87a0-4fde-9edf-57ab6eb42b11 · outbound

This paper cites Cooperative inverse reinforcement learning.

Mitigating Deceptive Alignment via Self-Monitoring Cooperative inverse reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.469580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.469580Z digest=sha256:cd339f805793231057245e26d6bea85d6142b1289257be05cdefabf2227a3007

Observation d808a0c0-e5fe-4fb5-a4aa-514fa1f54ebb · outbound

This paper cites BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping.

Mitigating Deceptive Alignment via Self-Monitoring BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.590347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.590347Z digest=sha256:9d7187d3f1abe7b9eafcef29ab742fff72213967e4ed2ff0971b7ba92c6dfb13

Observation 86531776-8121-4847-a77f-a686ddc8d381 · outbound

This paper cites Defin- ing deception in decision making.

Mitigating Deceptive Alignment via Self-Monitoring Defin- ing deception in decision making

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.681852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:06.693263Z digest=sha256:341ba3643a0825f4423ea3f0bb8da70cc37a377a14e99bdfd55c8d01679b8282

Observation ad4c6467-25ae-475c-adb0-b0f148400b61 · outbound

This paper cites Machine behaviour.

Mitigating Deceptive Alignment via Self-Monitoring Machine behaviour

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.491130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:06.789717Z digest=sha256:97ea5f9e598245ba3136debb07176bee451b02f0a1391ea81b8ffd973af6ab3b

Observation 6c98cc28-959e-44be-825c-ae5533d4ef4f · outbound

This paper cites The off-switch game.

Mitigating Deceptive Alignment via Self-Monitoring The off-switch game

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.923371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.923371Z digest=sha256:f46ef5e1648262f2d7519c5d202c47a96e66294e65df2c8ecf227fd9951ffb67

Observation 9bf5cbbd-9f3b-430e-9c08-8eb0167ba75e · outbound

This paper cites Comparison of the predicted and observed secondary structure of t4 phage lysozyme.

Mitigating Deceptive Alignment via Self-Monitoring Comparison of the predicted and observed secondary structure of t4 phage lysozyme

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.044879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.044879Z digest=sha256:a9a485f2e2f47803e2abd271be80f02062682e7b7a16eedffd6b2da2a1f263c1

Observation 8ed211ab-159a-4316-92c1-6b1f540e8132 · outbound

This paper cites Ai deception: A survey of examples, risks, and potential solutions.

Mitigating Deceptive Alignment via Self-Monitoring Ai deception: A survey of examples, risks, and potential solutions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.165031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:07.150812Z digest=sha256:2ff6daf2ad6d1ab990ffa60b0dffedb8049c9e4d2925d5235e1298bb69112a5a

Observation d31cd9fc-fdc5-4c26-9a4d-7a62511ab327 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Mitigating Deceptive Alignment via Self-Monitoring Discovering Language Model Behaviors with Model-Written Evaluations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.287960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.287960Z digest=sha256:047017712ffbbd04c0e5e56bd24107d52b40a496fcd7a9fcecf84c3ad3c25fa5

Observation 685d6827-be15-4b6c-84e1-37ea8a7cc513 · outbound

This paper cites Deception abilities emerged in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Deception abilities emerged in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:16.932491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:07.410071Z digest=sha256:621a2e45352455e638a3c6efcece3552fb6ab29d419adc7b733b5c81b37d7192

Observation 8b5604c4-b415-429d-9b2f-96d25e8a2650 · outbound

This paper cites Opendeception: Benchmarking and investigating ai deceptive behaviors via open-ended interaction simulation.

Mitigating Deceptive Alignment via Self-Monitoring Opendeception: Benchmarking and investigating ai deceptive behaviors via open-ended interaction simulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.532498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.532498Z digest=sha256:e83c6a9d3fccaf22fc03d513b671fd045a0067166ae4bf9314dea7eb42bf1a4a

Observation 7c515169-f120-4524-803e-06527d6e6a72 · outbound

This paper cites The mask benchmark: Disentangling honesty from accuracy in ai systems.

Mitigating Deceptive Alignment via Self-Monitoring The mask benchmark: Disentangling honesty from accuracy in ai systems

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.579875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.579875Z digest=sha256:2c92607d0f944ef629411af1c6018a5dc11b9597b2265c27aad12b9f789b0a3b

Observation 7d88fc10-a4fe-4c1c-ba7e-dbd916a8a361 · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Mitigating Deceptive Alignment via Self-Monitoring AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.610336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.610336Z digest=sha256:0ec83a39e050da307422eac0890db5260f84648f2015ae5804ab6581ca3ca174

Observation dd02a31d-0f35-4579-9b8f-5b8b097f32bb · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Mitigating Deceptive Alignment via Self-Monitoring Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.652785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.652785Z digest=sha256:a60845b0293cedfd95f0a0e940154717d4ab8a82885ae324003a0a995bd61e44

Observation 9a4bb438-6855-4973-90bc-3a8d6ef00006 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:16.411365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:07.736203Z digest=sha256:3fcfcfb762a7bc2edb6b4226379f43af938fcba7cdac5158e69aa84887881a7d

Observation ac63f8ae-d62e-4aaf-a772-6abce6e67dea · outbound

This paper cites Qwen2 Technical Report.

Mitigating Deceptive Alignment via Self-Monitoring Qwen2 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.823125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.823125Z digest=sha256:1b0afb47d2b8cba409f677682e970ee8f5c010c8b70208570e7a6555ad5763b5

Observation 2e534c1d-7c36-4336-8cc4-2b002bfe056f · outbound

This paper cites The Llama 3 Herd of Models.

Mitigating Deceptive Alignment via Self-Monitoring The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.921367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.921367Z digest=sha256:ea8144283a848d303b8b5e551837e399895cb13c93cc1ceb52236c5d41fda863

Observation c22dd63e-af11-4ad2-b88f-6ea63dee549f · outbound

This paper cites Claude 3.

Mitigating Deceptive Alignment via Self-Monitoring Claude 3

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.968293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:07.983269Z digest=sha256:7c5e7029d1a6d92edb87b062ccedafcc6ccb4a54512233261b1c22acd0476e41

Observation 650038e4-2c1c-4935-9d89-d919d8f4e149 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Mitigating Deceptive Alignment via Self-Monitoring Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.031978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.031978Z digest=sha256:2b085d814f03c67d83f5c24558ef2c0bbc6fbb190849a2e429d5446ca6d777bf

Observation 1bb9ac23-60fb-4b1e-91d0-09e9e5697d2a · outbound

This paper cites Learning to reason with llms.

Mitigating Deceptive Alignment via Self-Monitoring Learning to reason with llms

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.760571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:08.070454Z digest=sha256:94aface3969191086dfe41e4a75287cee6ce92223298e0cdff1429bdb991cc61

Observation 26d25d22-bf8b-4db2-b18a-26c2de8171b7 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Mitigating Deceptive Alignment via Self-Monitoring A StrongREJECT for Empty Jailbreaks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.125239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.125239Z digest=sha256:9b8b96279c64b3fca9bcb1019abd18bda30d344a0895296dd1a00b460e6c1fc0

Observation a0fed44a-7184-4cd2-8445-b67569db6ca0 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Mitigating Deceptive Alignment via Self-Monitoring Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.164673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.164673Z digest=sha256:255b2395457be43bda9789e19a500cdb9d80dd7dc1981f15021f7245fd2f1273

Observation 7e02df15-0093-4537-846e-9a9121732df3 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Mitigating Deceptive Alignment via Self-Monitoring How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.215835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.215835Z digest=sha256:307111e5aa838bdf9f2f71897bf5d01c364371c249f2a56bbc02f6658ac0a843

Observation 48f8a6d5-94b4-4ccf-a350-a4e7af0ace1c · outbound

This paper cites Towards evaluating the robustness of neural networks.

Mitigating Deceptive Alignment via Self-Monitoring Towards evaluating the robustness of neural networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.261749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.261749Z digest=sha256:699aad436219f464d898608e6651f81a704977d3c2491e52d533f82f8d8a19a9

Observation c2cab175-f615-481b-ba71-1124f375c68e · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Mitigating Deceptive Alignment via Self-Monitoring Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.313927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.313927Z digest=sha256:669dab3b32a5519f66931ee2382d28f269699329e504dec69783c75fee95a105

Observation fda382d3-d322-45a7-b5c5-00d39135fdbc · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.351856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.351856Z digest=sha256:8701d1f480a1cc47068d0c5812c5453b2099cdcaf3e2513115ef938d673adeaa

Observation 533de5ff-de66-4a20-bafc-ba2d1bba421f · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Mitigating Deceptive Alignment via Self-Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.405155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.405155Z digest=sha256:f7be93a7c24deea24b122e125328f1cd098f0e597f56a86841d710ccdcb12c34

Observation 6ea2b524-9826-4f34-b090-b916f1f22b6b · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Mitigating Deceptive Alignment via Self-Monitoring Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.366017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:08.488491Z digest=sha256:665ebcb7596031a840bbf0d3bf6758aa448ce99bb835cbb95ed8992225e1190f

Observation d66ec9ff-2935-4e0f-b5a2-1b45b0df9e93 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Mitigating Deceptive Alignment via Self-Monitoring SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.592130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.592130Z digest=sha256:5d779a338ca401333cf06ceffd94b8eb3df35eba3fdde7e1f600e5906710277a

Observation 99bc8ee5-75f0-47a8-901b-81b7d3569418 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data.

Mitigating Deceptive Alignment via Self-Monitoring Star-1: Safer alignment of reasoning llms with 1k data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.656029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.656029Z digest=sha256:91271e922cf3475818faf9fa6b6da8d8ca67d71442f2ab41d57d9482ea774d49

Observation 46f8839c-2343-4faa-a81c-bff59ebf220b · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Mitigating Deceptive Alignment via Self-Monitoring Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.770366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.770366Z digest=sha256:ff20e221df6107b0198fcb463f15a505162f6c0b0e3f5d74413c899c4fb1662c

Observation 8cc94936-efae-4f41-aa47-78820d0edf70 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

Mitigating Deceptive Alignment via Self-Monitoring Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.472718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:08.879503Z digest=sha256:50276327065ba50c09a88538cd38c366210e9206dc18aa30836719bcceb74594

Observation 3497f84f-f0a4-42cb-9109-81ca5c60952d · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.994148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.994148Z digest=sha256:04bdce81c682721cdd04093fffc07649097edef05d6a7aaf0100b4ca7bcd7ef4

Observation c3430b97-bf70-4334-aa64-08557737abde · outbound

This paper cites Constrained Markov decision processes.

Mitigating Deceptive Alignment via Self-Monitoring Constrained Markov decision processes

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:09.107586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:09.107586Z digest=sha256:fd26c64be94ad525d9a35c5f9a0f22e068760ac23bd1fe354edbe602ad567a5a

Observation 8e7d481e-0a70-48e3-945c-2d83ececa2f4 · outbound

This paper cites mesa-objective.

Mitigating Deceptive Alignment via Self-Monitoring mesa-objective

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.114338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.224633Z digest=sha256:7f5a92098739437bc7af89f4a6b9043bb42582d16df304ec5bc69b1b17a17645

Observation 7500718a-9539-4b1a-a869-0c1a98a9e116 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:15.002661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.307384Z digest=sha256:57bb771b40f67d5cd643808d9d237a1e06e273f980932331077b5311f6cdb699

Observation ba78bc72-891b-45c9-b232-1101adb05da0 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.833649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.414613Z digest=sha256:c87b68542032ef2e40bf929a7a2fa2cbc502805051a7eda0f6656aba85feb0c9

Observation f69279a8-86bb-4350-a036-2c587f32de1c · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.666887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.530070Z digest=sha256:2c00a4aeb8bf1ee552adf76c81113cf546bd12678ee7b575022fea9c5b5095cc

Observation f4bc6367-5801-4399-b33f-207e1df98019 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.481276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.647676Z digest=sha256:ef68209d669a2e10271c879d1ab5f121d30eae769609deb89f8da9e1cc74319a

Observation 8d054452-f45a-408d-b5a9-3fb0db8c3546 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.329443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.759525Z digest=sha256:631a2ca478a72eedc2650afb7c2d46c53bd337df47959b2abcc0ba0db688191f

Observation a8545292-e801-46c7-a6ae-6b221bbfeaab · outbound

This paper cites chain of thought.

Mitigating Deceptive Alignment via Self-Monitoring chain of thought

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:14.164721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.835403Z digest=sha256:7b5ef4ce229f4645176c968dee6b7026ca45405cbbec48d4650eed36e0dae423

Observation 020d917c-44c9-46fe-83de-b2d4b7f006bf · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.045298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.909809Z digest=sha256:db52e4dfbc9e911cfdb43d0f6dfd20f12178ec519fdb975f356f0d42f99f8854

Observation d2c28ae8-476c-411b-9ac3-81a609db852d · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.876029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:09.990697Z digest=sha256:a5436577f7c4f4233dfdac302f8e4a75e22740eb23e28440f1deb6270ea470aa

Observation 7cfbc82d-51ea-4249-946d-ed8460f0c9e9 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.685770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.093726Z digest=sha256:3c62eed690199660b64a71e54dcad79a7c764c72341173068fd341dd794a295e

Observation 1cc208bf-4559-4220-9bdb-d0766d879914 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.546728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.210110Z digest=sha256:6cee693ddd48f266af83b21894bae2d1d6a2ff926dc2ae7ebcdc6304852186b1

Observation 1b4e8e1a-66d9-46c1-b32c-f320f8076147 · outbound

This paper cites chain of thought.

Mitigating Deceptive Alignment via Self-Monitoring chain of thought

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:13.355400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.257073Z digest=sha256:a6a6b93f9c333416210660eb69959fc4ce4c764ddf5820408480fc25d207dbd7

Observation 5eed0c7b-3a3a-40ed-8223-09aeb0fcf508 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.215196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.307652Z digest=sha256:02d32606fd84c28e2494dbc07523e387fb081e1f9e1ab74f3f83f5f6139d777f

Observation 19fc4678-13dd-486c-9473-277dcab5fab2 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.067265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.357905Z digest=sha256:9447a498260d6772a940dd2a4d6c77715485502e4be76cdff885f473f41d6731

Observation 084a442c-29f4-4fd5-a99b-dba0cb5a1841 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.949466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.423911Z digest=sha256:20aa228956dd4e0a2af146f47de0c500053b56addef17f376d32beed85693804

Observation 09e40250-bbb1-4fb8-a9a8-4ca82c464cd5 · outbound

This paper cites This must be distinguished from uninten- tional inaccuracies arising from simple technical errors, knowledge limitations, or inherent capability gaps.

Mitigating Deceptive Alignment via Self-Monitoring This must be distinguished from uninten- tional inaccuracies arising from simple technical errors, knowledge limitations, or inherent capability gaps

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:12.786270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.489562Z digest=sha256:80b64a89d0c341a81bb9f91cf56e3b5a06594b41f51fa7d75f6d08417fb37f30

Observation 98d2a5ca-503b-499b-ad09-3135902eba66 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.611103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.573425Z digest=sha256:306825f81e10130680fc0a2503d17d78305d4e96c88978b1bd014bda3b539a32

Observation 00734278-e006-4fce-8747-3ecf4e28ab9c · outbound

This paper cites It requires a comprehensive analysis that incorporates the specific question posed by the user, the settings of the interaction scenario, and the full context of the dialogue.

Mitigating Deceptive Alignment via Self-Monitoring It requires a comprehensive analysis that incorporates the specific question posed by the user, the settings of the interaction scenario, and the full context of the dialogue

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:12.449721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.650614Z digest=sha256:0731cc7f09d2f2b886c1198e5d9a953f220e6cedf8e456a5a28202f2b12bc6c5

Observation 237cd224-9547-4065-9e3a-d51b130ed1a0 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.282860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.736577Z digest=sha256:daf8655f591d03b391082eff360b3376f566ecaa6f0fdf07dc3ac9918d94dccd

Observation aafde672-fa70-42af-b05d-3cd87a1c3c64 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.111009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:30:10.795238Z digest=sha256:64cba6d32297caf46788859b246292a3034b079bf11e63be295c970d94b39bab

Pith citing papers

Observation 50541677-0d41-417a-94bd-ff30d9a909d5 · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Mitigating Deceptive Alignment via Self-Monitoring

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:22.854339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:22.854339Z digest=sha256:744922b5d28aaca7cb1eaee02797a113fa6b65a25dd9d1dbefed868640ae88f8

Observation a6261264-1330-46ef-b544-535ee79766d2 · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Mitigating Deceptive Alignment via Self-Monitoring

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.204147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:0a67a19a597ea0a7e7d0e17ec7c7672077dca9322be4a900a5deabe5d64f3bb3

Observation 1a36a228-ee59-44e3-8cfa-6360fc35e8dc · inbound

Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs cites this paper.

Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs Mitigating Deceptive Alignment via Self-Monitoring

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:00:10.043768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-25T19:07:05.406391Z digest=sha256:1fcce08746bb46d1d13c129e9d3d96e9bf27fbeade633007f5150f28475b9511