Pith. sign in

Paper Citation Record · LEDGER

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 24 inbound Pith citation observations for arXiv:2508.17511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17511 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:57:10.334753Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T16:13:13.602057Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7b365787-74ad-4a0e-a2ff-bf766966df4f · outbound

This paper cites write newline.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.274752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.274752Z digest=sha256:c87e09297324e54a211be5df88dd69dbc93d9dcc9573ec8481b398cf68082a35

Observation 1361b49c-213a-448a-89a1-b75ab9d6cb4d · outbound

This paper cites Program Synthesis with Large Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.350528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.350528Z digest=sha256:498834da35a241463a88a4dc8fc2199c2428fdd9661744787af6e41a202f3090

Observation 06bab222-0455-494b-9d93-a2aa1234a229 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.434931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.434931Z digest=sha256:316c43f6c5ebbb9015c43f8707230cd35334a0ef9391d900e79b270e88a0029a

Observation e33121b4-fdb1-4485-a060-897e7dc12f0c · outbound

This paper cites Tell me about yourself: LLMs are aware of their learned behaviors.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Tell me about yourself: LLMs are aware of their learned behaviors

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.514379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.514379Z digest=sha256:ecb78b402ca609756e46aa27ea92834db9f1b8ce80270508c1628afbe1858dd8

Observation f28fba28-da30-43a9-861a-2120459b5764 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025 b.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025 b

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.604117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.604117Z digest=sha256:14a7d995d70e6451ad9ff3cb7a6d4bf6f00ab576fb5c03c52c4e520b75c67ca9

Observation 6aa00b91-28f3-469f-ae21-03a5cdd1d5fb · outbound

This paper cites Demonstrating specification gaming in reasoning models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Demonstrating specification gaming in reasoning models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.678736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.678736Z digest=sha256:a8e190d2013e46a5d2f327474a5792dd7b2c8e9c06aa836b0463ea51fa771f2c

Observation 98d8e0f3-96d2-4ea4-99e9-64d3198500ee · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.784533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.784533Z digest=sha256:c7f06cfaa0d826fbf2baf9d074fc68955c627d6199857576d60adfc476710f48

Observation d127ecec-5b7a-47ac-9555-212bc4cadb32 · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.932075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.932075Z digest=sha256:e0271cf66f06d97d826510c3cbf97ff0369e70f5b97bcaaf96be5915f7ed6adb

Observation 890c0684-6695-4f36-94b4-7a6787d561d7 · outbound

This paper cites Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.044929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.044929Z digest=sha256:db51b858d7d32b51f2d4a95da5e5a0345e4d17b1ecfe0577b2b0b23e9c25251b

Observation 78da8815-e0ee-4ce5-bc17-5d4387958ff1 · outbound

This paper cites Subliminal Learning: Language models transmit behavioral traits via hidden signals in data.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Subliminal Learning: Language models transmit behavioral traits via hidden signals in data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.146810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.146810Z digest=sha256:94ecf1688b8ca261589cb4f8cb2cd3114b1c3cce5e23c0ab73b5c18b664d96dc

Observation 7dfb1ece-62a5-4ac4-8c64-64ee1d7fcfa3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.253618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.253618Z digest=sha256:3f27ce562baa0980b025ea6f07b1656554431f2ce365063dcb615c4f481ace5a

Observation 70f59ac2-da0b-4613-aa8f-bdf632c4fa36 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.297166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.297166Z digest=sha256:6ab18adcca561b40fd7676d7686984bc03cd40884eab21826839b47e5c0e5e4a

Observation 5956c521-90a0-4cf8-b52d-4688324df1bb · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.311545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.311545Z digest=sha256:a62fcdd67e2d165002c461a51e12c7bb5d935020107512c042ac2ca181a792a2

Observation 784a3ef4-a4c9-4a12-a593-bae8ce0e12b1 · outbound

This paper cites Unsloth, 2023.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unsloth, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.416810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.394755Z digest=sha256:3418356db0fdf698987ab02e8fe1f8ca8f83d15988ee3546cdc34679c0f37c1e

Observation aa0012d9-36fa-4c54-acfb-7ad50297b706 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.534766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.534766Z digest=sha256:85cc8bb6307d59f4ec393b3bc654368d49aa2da04c6ad85e695844b7208efab2

Observation eda3174a-4a4c-4a1b-9f3b-5ca34c117027 · outbound

This paper cites Training on documents about reward hacking induces reward hacking, 2024.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Training on documents about reward hacking induces reward hacking, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.371924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.663990Z digest=sha256:8c1bcbe78c4bb86c7804fd6f3aa21638954bd400a9c0c50ce526291ae254df0b

Observation 621dfae9-4717-4e66-b334-eec7140e44da · outbound

This paper cites Model organisms of misalignment: The case for a new pillar of alignment research, 2023.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Model organisms of misalignment: The case for a new pillar of alignment research, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.353791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.765494Z digest=sha256:d8b0259f764cda774ebef42c73c2bd163b1004c98ebe7d947622013a2b4387b9

Observation 70e3acff-d50d-4d41-8bde-f83b37011a63 · outbound

This paper cites Auditing language models for hidden objectives.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Auditing language models for hidden objectives

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.878310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.878310Z digest=sha256:205f8d4c8f947cff10085697eb781179e902b15b09c3a742bfe617386add9861

Observation 863f3f84-5cf5-4e51-a133-e44a69dac6f4 · outbound

This paper cites Recent frontier models are reward hacking.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Recent frontier models are reward hacking

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.337301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.917897Z digest=sha256:3918a126453b3fdef4dca92c1e8bd6df4f1896833d2848557b3919370ee07430

Observation 8d0e2e04-7f0a-4ff2-855d-d82743908a90 · outbound

This paper cites Reward hacking behavior can generalize across tasks.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Reward hacking behavior can generalize across tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.072971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.072971Z digest=sha256:4ac282722adbf7e7584064e50f908fed4583c63a1db0902e69941445496dade2

Observation 5d75d60e-1a81-4e36-8a1f-44b29c5895e0 · outbound

This paper cites Toward understanding and preventing misalignment generalization, 2025.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Toward understanding and preventing misalignment generalization, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.296022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T16:57:09.228351Z digest=sha256:558762e475a4d963570e378c1fb4a285faf5f648d96d8bc7ac7073158b12533c

Observation 7e2ccd9f-447a-4132-9eec-4fb7a6863344 · outbound

This paper cites Sycophancy in GPT-4o : what happened and what we're doing about it.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Sycophancy in GPT-4o : what happened and what we're doing about it

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.260052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T16:57:09.317538Z digest=sha256:54d91b23b1d17754498247af2a84576246363c17d4cdc8607ae1e7b6920748b0

Observation a6a0e824-aa5d-4f24-a3af-6aa1b6678cc2 · outbound

This paper cites Generalizing verifiable instruction following, 2025.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Generalizing verifiable instruction following, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.412774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.412774Z digest=sha256:92e63202f03bfc85f7ed9fa5db2dbb991a6b38016b644e77299a3a5944c518e9

Observation 2a5b2435-3baf-4cb5-982e-b69d4fdaeaf2 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Towards Understanding Sycophancy in Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.417680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.417680Z digest=sha256:bbd302ed56619ca0cc8ebdbbf2cea41c27d1e483914491725afd3fe114192575

Observation 6e307ebc-94ba-4061-acc0-74ecd3b48a5a · outbound

This paper cites Defining and Characterizing Reward Hacking.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Defining and Characterizing Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.545711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.545711Z digest=sha256:5f457535932549d6357b19646ec36219938a51eb3fa25b5a2d9e42779fab5381

Observation 7b69b142-7d9a-480d-8ccf-740e5d729a47 · outbound

This paper cites Hashimoto.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Hashimoto

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.649343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.649343Z digest=sha256:40e500bcb9a4ef4ca1f46fe8679a0845306e9a4698acfcc842e20a6eb2a36852

Observation b5cb4994-d88d-49f1-b69a-f41119e1df3b · outbound

This paper cites Model Organisms for Emergent Misalignment.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Model Organisms for Emergent Misalignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.784760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.784760Z digest=sha256:a0756d5dff6601b1400f6428eeacbc77daeba93c88131db176db8e2953d43187

Observation 981e1df8-dff4-4d22-b14c-f8a33ad5cc05 · outbound

This paper cites Chi, Samuel Miserendino, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Chi, Samuel Miserendino, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.862801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.862801Z digest=sha256:24c4a50a28761647daaa1fe61b5fff5c7859277c1c477a3a67900fe89bdaa822

Observation 7eed511f-09fc-4d4d-b5e0-a114c3a26595 · outbound

This paper cites @esa (Ref.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs @esa (Ref

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.995780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.995780Z digest=sha256:1a475c9ed0ed3c65b022921eed7e2ff13452c9e5331529ee06e02bdb7c025b42

Observation cdd9a63c-5706-4460-9fa7-36d3e3b8b84d · outbound

This paper cites an unresolved cited work.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:10.144751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:10.144751Z digest=sha256:c0171b5e15130ba5a9773a0ea4fde5ecf7e7072dd027e46eb60ec327b90b756c

Observation fdac59e6-b602-4065-8c97-ed4869c5008e · outbound

This paper cites an unresolved cited work.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:10.334753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:10.334753Z digest=sha256:6f40dae6e86c132856b9f44175bfd68ff398a1b66cb17012903691a1dd06536d

Pith citing papers

Observation ce34dc64-f522-4603-bdf4-1c2dc13ab2bc · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:9248a885fb2484a6baf5741be6da9a73ecaba51cccbff6c0cfe4f00c4b1fbf70

Observation 7f3ed349-c51f-4c6e-b603-6a3523d6430b · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:37.804215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T10:41:14.970156Z digest=sha256:f0e48fd42cd20d80ade761baf6811d784e5a148af6b59228089687d0b386ed96

Observation 3eda4d5e-7340-4e44-a24d-f3be0c6e96a5 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T16:13:13.602057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:13:13.602057Z digest=sha256:eb20e182ee5eda558c95c864db5291539c80378f2cb766ac83744f2169842af8

Observation 3453397f-e49a-4967-9f2c-70bce11750c7 · inbound

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation cites this paper.

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:10.353589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T06:38:32.413172Z digest=sha256:0bd6f7b032ca70888bacf5c7445b481f68c883cd7c65b208d2bf409c073c8959

Observation bfe7a9a8-26bf-41f3-9b9f-dfa4133a547a · inbound

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation cites this paper.

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.210533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T09:56:48.019404Z digest=sha256:c11a926c6ef0c4a5cbe51b5c942f7d046aa5aec4817189a417df01c22b07842b

Observation f5231124-49b6-4f7c-aada-3723f8f144b2 · inbound

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use cites this paper.

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.643957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T15:38:41.821264Z digest=sha256:54053c2fc3477720ed652566eda4eccadc5b8ad572359e6c373576c2a2c84ddd

Observation 95a83a75-de8e-4aa4-af65-8619b902b83c · inbound

Overtrained, Not Misaligned cites this paper.

Overtrained, Not Misaligned School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:47:26.302602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T06:45:52.544674Z digest=sha256:f633441bd2602c66132852f0c2f7d10a8d55522a9c329ad33c2d552334f18e6c

Observation f0aa3db9-adac-4d18-b1a8-16176f80e41a · inbound

Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning cites this paper.

Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:18:18.493663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T13:13:51.081597Z digest=sha256:ca4bc1e91a1f23c20f4be28110c91cbf1671d1346aa37d3e2d22df25eb61c3ff

Observation db44bc86-3e8f-4461-9f2f-01526289e32f · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.563392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:cc018c108826690220e6e9f991320237ec270163b5d6cd339e2ee1d099ad6cab

Observation 2cbb4fd0-967f-45d3-b779-fed347e46efd · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.710993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:2e002f9c5d72a0dfd5d87d97e795f5de3b113cac2c4616ded8ad65a713435df7

Observation de3ed3ac-8c46-425b-94cb-335da2faf258 · inbound

Relational Intervention During Functional Collapse in Large Language Models: A Lexical-Statistical Ablation and a Structure x Register Factorial cites this paper.

Relational Intervention During Functional Collapse in Large Language Models: A Lexical-Statistical Ablation and a Structure x Register Factorial School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.083579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:43:14.182982Z digest=sha256:ea7c72a01ca3f3a60b86d60d8288f37df3ce82469fe1a84d55afd3ba3a95ea7f

Observation 30d87d04-b386-42da-8b9a-857b0cb689d4 · inbound

Consistency Training Can Entrench Misalignment cites this paper.

Consistency Training Can Entrench Misalignment School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:28.625276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:07:31.153337Z digest=sha256:83f98a12433407b3b298a8fdb0996fec3fb4de66085fcd07bf47f979b8d90b46

Observation c1cabe89-00eb-4dfe-974d-8ca33f71aff2 · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:36:59.372429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:07:17.014002Z digest=sha256:2595e86f4d2241ae32a705f5692f6c87fc695c8af667bfd5a02a1509539fa51f

Observation 4d4e8120-96da-4e2a-add1-443a7f1911b3 · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T12:19:47.379539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:19:47.379539Z digest=sha256:30fb448fe0042e3d76d337c45d34b4a750e2d116ade417cb581d972aad17e555

Observation c7cd1825-11ed-42a8-8992-dd38d09f7511 · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.086544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:8f0d9c9df66816501233a473b29ef9d38d98e4bada4c88b58f9286c5f5d60c74

Observation c0078213-a2f5-42f5-971f-16695aa02b9c · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.443095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:ba4adde35eadd0e04e6dcb04df1267149989be4ef9c065d558392b6024e464b8

Observation e6ccf1ec-ceeb-440b-9232-a261d8878e3b · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:31:02.725240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:54f73dbe35ef6296f4e546ecbe68f9fe25dcba4f39e52b39cb376935a1545ee6

Observation 1487e9c4-3997-4628-b974-4b68a63739e6 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.116580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:e2ef5ab7eb635546862488fe152094a983899621fb36271788bdf781ffa3af7e

Observation 9dd787f8-3947-4110-bbbf-511e45f6b6fb · inbound

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment cites this paper.

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.786188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T02:39:07.685462Z digest=sha256:620760ad863d806c818d1d9d1d5678cee0f571c4159d41e4faaeb618af7f4550

Observation e9ab84d1-fcd4-416b-8688-139ea89f0c48 · inbound

Reinforcement Learning Towards Broadly and Persistently Beneficial Models cites this paper.

Reinforcement Learning Towards Broadly and Persistently Beneficial Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T09:09:16.527440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T07:51:13.283619Z digest=sha256:2885f737be3aced2fa5adea666f24124ed9377dbd3f856dc1121d232972cad1a

Observation 4a934e37-1531-4e29-92d3-74f614d0c338 · inbound

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? cites this paper.

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T00:42:24.432562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:42:24.432562Z digest=sha256:2c3556af1ba4732a81de7a47c0c036902d41199be9ac568bc8e5c191dc2c1374

Observation 8f3512b3-1275-485f-9d8d-f36b03c8d1e9 · inbound

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs cites this paper.

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:51:18.302638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:51:18.302638Z digest=sha256:beb9dfee6d02f8e695b43470e95602fe325708aa1f720a194f64994148d8a88d

Observation 05ec6505-e1fd-4ef9-a787-9d93054c8cf3 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:22.402195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:22.402195Z digest=sha256:1584c955fab83418c71b01bf3847c082c8d11e9a645001a6d804580ac33f2cf8

Observation bb5ea07f-7d12-4c48-8ca3-f5cd63365fc3 · inbound

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models cites this paper.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.908101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.908101Z digest=sha256:391e68ce66325e8e2e00a3ea56e43870fd6a8d9076a07990362d004fec198848