Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:49:58.910725Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2508.04216.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:49:58.910725Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T12:22:06.635622Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T12:23:24.240761Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 53475a56-da37-4582-bb18-dfc7b62a9ad8 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction , " * write output.state after.block = add.period write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21285f0-4606-462a-8b62-e26e1c89cb44 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7953bb5a-bb0d-4a5f-bbc0-e78ea0e11043 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Concrete Problems in AI Safety
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426f3b28-a921-476d-88da-4a3fe872c68c · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 870055df-2744-4273-8716-6b2db294e08f · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction E.; Hume, T.; Carter, S.; Henighan, T.; and Olah, C
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56d4c56-8dea-4b89-a9fe-5ff1ff83ac11 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7532f3ab-76f8-46a2-9748-bac3186db987 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4fed5e4-9b22-432b-9197-85b9a818c276 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48843abe-3a9f-443d-aca0-0ce19980c0d0 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c175fe98-219e-43b5-be7c-dcaddc4ccf1d · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b720003-3c15-4315-87ef-c01af2f16297 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff348be9-27de-427a-8ce6-a0f67c93a558 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Inverse Reward Design
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993c3d33-30a8-4051-9fcd-da0bce7e8267 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02c988b7-9dd4-47ae-bba3-607f415acebc · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Let's Verify Step by Step
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1a11d3-de30-4c11-9006-604f787b15e8 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151a3599-716c-4f98-ad09-f762cebc36b2 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2786640-1018-43ad-8f4e-99a900c9a45c · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41967675-9e22-4ff0-9d31-615883143be6 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f01d6fc3-fd5d-46be-9266-4d1dc61a95e1 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5fc8f8d2-3bc5-4065-acc1-1db253b092fd · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Feedback Loops With Language Models Drive In-Context Reward Hacking
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c35825-02bd-434b-9fd9-eacef2c06611 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b599c292-0fe5-4857-8b69-41716d6dfa30 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11feddce-4347-4d37-8522-ca4ea2a71324 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c81a857b-41ce-4035-aa10-4432d0b99b1f · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Self-critiquing models for assisting human evaluators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aef55fa-a437-473b-bcd4-c93e13044071 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3abde802-218f-44d7-ba48-e00c9527e919 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d45b03e4-3e88-475d-a03f-6465fd3d4c36 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca4aa02-c030-4f7d-9025-fd91360e5b06 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bf3754-bb08-44c9-8096-667569423794 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8048e73-95ee-4f11-8082-d620ed4e4900 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction V.; Lee, J.; Xu, K.; and Kumar, A
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94254e3b-f5aa-4a07-87f9-b90e7bc10d02 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06cff769-dda9-49ac-877f-d6f01645f5de · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea48494e-966a-4450-9252-0561cecf8267 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction L.; McDougall, C.; MacDiarmid, M.; Freeman, C
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa39720-4336-4778-996c-69b162fd43dc · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Solving math word problems with process- and outcome-based feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e4ba43-a24d-4141-b814-b64088f259ae · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5210ed7d-8a0c-4e7d-8e2e-8a4285517db3 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction V.; and Zhou, D
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bac4398b-b0d7-4bdf-9178-1545702fd1ef · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a68159bb-9a83-431f-834f-94b0921806ab · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99f1c65a-14e5-4d15-8569-c13c054177e4 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941c3b9b-e322-4766-ac12-5d0c6c6af4b9 · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction L.; Cao, Y.; and Narasimhan, K
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87d5d540-9d51-4cb6-85c3-5b5292754f9a · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Automatic Chain of Thought Prompting in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb593af-902e-48f0-b2d3-c34250c53f0b · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23122d00-eb26-4d5e-92d0-df77482177dd · outbound
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction A Survey of Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74308e21-8960-496a-9d51-81d753f7543e · inbound
Factored Causal Representation Learning for Robust Reward Modeling in RLHF Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 90bf14c4-aaac-4d72-a5a3-94d7926aa0ef · inbound
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.