Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:57:10.334753Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 24 inbound Pith citation observations for arXiv:2508.17511.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:57:10.334753Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T16:13:13.602057Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 7b365787-74ad-4a0e-a2ff-bf766966df4f · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1361b49c-213a-448a-89a1-b75ab9d6cb4d · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bab222-0455-494b-9d93-a2aa1234a229 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33121b4-fdb1-4485-a060-897e7dc12f0c · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Tell me about yourself: LLMs are aware of their learned behaviors
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28fba28-da30-43a9-861a-2120459b5764 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025 b
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa00b91-28f3-469f-ae21-03a5cdd1d5fb · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Demonstrating specification gaming in reasoning models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d8e0f3-96d2-4ea4-99e9-64d3198500ee · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d127ecec-5b7a-47ac-9555-212bc4cadb32 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890c0684-6695-4f36-94b4-7a6787d561d7 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78da8815-e0ee-4ce5-bc17-5d4387958ff1 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfb1ece-62a5-4ac4-8c64-64ee1d7fcfa3 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70f59ac2-da0b-4613-aa8f-bdf632c4fa36 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5956c521-90a0-4cf8-b52d-4688324df1bb · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 784a3ef4-a4c9-4a12-a593-bae8ce0e12b1 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unsloth, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa0012d9-36fa-4c54-acfb-7ad50297b706 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs LoRA: Low-Rank Adaptation of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eda3174a-4a4c-4a1b-9f3b-5ca34c117027 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Training on documents about reward hacking induces reward hacking, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 621dfae9-4717-4e66-b334-eec7140e44da · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Model organisms of misalignment: The case for a new pillar of alignment research, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 70e3acff-d50d-4d41-8bde-f83b37011a63 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 863f3f84-5cf5-4e51-a133-e44a69dac6f4 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Recent frontier models are reward hacking
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d0e2e04-7f0a-4ff2-855d-d82743908a90 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Reward hacking behavior can generalize across tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d75d60e-1a81-4e36-8a1f-44b29c5895e0 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Toward understanding and preventing misalignment generalization, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e2ccd9f-447a-4132-9eec-4fb7a6863344 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Sycophancy in GPT-4o : what happened and what we're doing about it
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6a0e824-aa5d-4f24-a3af-6aa1b6678cc2 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Generalizing verifiable instruction following, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5b2435-3baf-4cb5-982e-b69d4fdaeaf2 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Towards Understanding Sycophancy in Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e307ebc-94ba-4061-acc0-74ecd3b48a5a · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Defining and Characterizing Reward Hacking
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b69b142-7d9a-480d-8ccf-740e5d729a47 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Hashimoto
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5cb4994-d88d-49f1-b69a-f41119e1df3b · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Model Organisms for Emergent Misalignment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981e1df8-dff4-4d22-b14c-f8a33ad5cc05 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Chi, Samuel Miserendino, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eed511f-09fc-4d4d-b5e0-a114c3a26595 · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs @esa (Ref
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd9a63c-5706-4460-9fa7-36d3e3b8b84d · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdac59e6-b602-4065-8c97-ed4869c5008e · outbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce34dc64-f522-4603-bdf4-1c2dc13ab2bc · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3ed349-c51f-4c6e-b603-6a3523d6430b · inbound
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3eda4d5e-7340-4e44-a24d-f3be0c6e96a5 · inbound
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3453397f-e49a-4967-9f2c-70bce11750c7 · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bfe7a9a8-26bf-41f3-9b9f-dfa4133a547a · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f5231124-49b6-4f7c-aada-3723f8f144b2 · inbound
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95a83a75-de8e-4aa4-af65-8619b902b83c · inbound
Overtrained, Not Misaligned School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0aa3db9-adac-4d18-b1a8-16176f80e41a · inbound
Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db44bc86-3e8f-4461-9f2f-01526289e32f · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2cbb4fd0-967f-45d3-b779-fed347e46efd · inbound
Understanding Goal Generalisation in Sequential Reinforcement Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de3ed3ac-8c46-425b-94cb-335da2faf258 · inbound
Relational Intervention During Functional Collapse in Large Language Models: A Lexical-Statistical Ablation and a Structure x Register Factorial School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30d87d04-b386-42da-8b9a-857b0cb689d4 · inbound
Consistency Training Can Entrench Misalignment School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c1cabe89-00eb-4dfe-974d-8ca33f71aff2 · inbound
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d4e8120-96da-4e2a-add1-443a7f1911b3 · inbound
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7cd1825-11ed-42a8-8992-dd38d09f7511 · inbound
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0078213-a2f5-42f5-971f-16695aa02b9c · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6ccf1ec-ceeb-440b-9232-a261d8878e3b · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1487e9c4-3997-4628-b974-4b68a63739e6 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 201
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9dd787f8-3947-4110-bbbf-511e45f6b6fb · inbound
Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9ab84d1-fcd4-416b-8688-139ea89f0c48 · inbound
Reinforcement Learning Towards Broadly and Persistently Beneficial Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a934e37-1531-4e29-92d3-74f614d0c338 · inbound
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3512b3-1275-485f-9d8d-f36b03c8d1e9 · inbound
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ec6505-e1fd-4ef9-a787-9d93054c8cf3 · inbound
Emergent Misalignment Recruits a Pre-existing Persona Subspace School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 207
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5ea07f-7d12-4c48-8ca3-f5cd63365fc3 · inbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.