Pith. sign in

Paper Citation Record · LEDGER

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

As of 22 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 7 inbound Pith citation observations for arXiv:2506.22777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22777 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:03:55.246559Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:51:52.262506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:34:47.983586Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f21d7222-a4fa-4ed0-83a0-10f82d45ea24 · outbound

This paper cites Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:50.441783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:50.441783Z digest=sha256:74c766c7b5f536be23821e5a825504c98e1764186a6354f2eeb97607be795c16

Observation ff74f8e3-c968-4801-8c60-1943027adc10 · outbound

This paper cites Claude 3.7 sonnet system card, 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Claude 3.7 sonnet system card, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.747302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:50.497422Z digest=sha256:0f78e3002746bfa0b68ebc8f166edb1457a72c2a37f3f22db4eaa5e4ca132023

Observation 302fb93a-fc55-4419-bffc-fe5132b7ef40 · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning System card: Claude opus 4 & claude sonnet 4, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.576798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:50.661471Z digest=sha256:a4ba8c5ea13f123c6198eb20f378cc386d7f1cbc2e7212bb5580a3683083650c

Observation 2ad3bb70-9718-4362-9b7b-5b5f34ad5512 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:50.831268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:50.831268Z digest=sha256:821d1485e965975abc57375884bcc80ba7f80c7f58e1ad21b7d32b32d01e0bec

Observation 239f0235-1f8c-40fb-b310-915a321296d4 · outbound

This paper cites Do models say what they learn?, 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do models say what they learn?, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.427022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:50.892798Z digest=sha256:3bc1566bac27cd73ce4b06fe31cd2a107345985a2ccfe3448487891cf931e91d

Observation c841f2e4-0213-4768-924c-436d2805774b · outbound

This paper cites CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring, May 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring, May 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.189283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:51.084035Z digest=sha256:29fb3fc266d3664194268243468a77e83ddd98ca67e84fb4d71bc0fe1f5efb48

Observation a3b04bbd-4df7-4467-88d2-39c5d647e735 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.211271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.211271Z digest=sha256:07d57c74b78c8e0fde07d50f29c574a8e905ffc315abbfbecca4dc679653ac13

Observation 616b1434-3fc4-4f52-b899-1ce6c2dff02c · outbound

This paper cites C., Macar, U., Nanda, N., and Conmy, A.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning C., Macar, U., Nanda, N., and Conmy, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.996924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:51.342380Z digest=sha256:748cd4e44aa0a73361bd9753f20b0f9db8ce028782c8bc1c2774c33829726c84

Observation 1bde6da0-f4f3-4a8c-9f28-1afffd67c247 · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Progress on Scalable Oversight for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.583235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.583235Z digest=sha256:cbe050808c75bdf46e4e4b6960aaedff5b5f35cffd82fdf77964974302bb4a67

Observation 54d38c93-7ef8-4d85-8794-6adf0cf2c45a · outbound

This paper cites Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.721948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.721948Z digest=sha256:617ab984c374f203f8977d4cf412249c601d41cc5e4652214cbf01dcde9a5513

Observation 34c38ac9-afcd-4cee-aaa6-c721edec54f4 · outbound

This paper cites Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.842152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.842152Z digest=sha256:6840ded8f21c00cf2c6978be157f3a32b3ad021aa0055ef49b5f273f0148357f

Observation 2170fee5-fc74-4209-a415-4a639f0c0b89 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.958119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.958119Z digest=sha256:100157eaf3619153000ac2b08cc0c7fe345c59990d9171bfa3a7ba08884eb25c

Observation dc59f3ac-4435-40e3-8787-ad58d3430371 · outbound

This paper cites Are DeepSeek R1 And Other Reasoning Models More Faithful?.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Are DeepSeek R1 And Other Reasoning Models More Faithful?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.114166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.114166Z digest=sha256:c25ba6c0e1f9212ea310fdc67e696d49eb095ac5d495e461eb1fb4368b13ae52

Observation 246b518d-1ab1-4735-b78f-3311d85ac3f2 · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.256909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.256909Z digest=sha256:da14e0a02e3801cdd2595be8043f3ee1e24afe45cc473d521ee03889493d42d2

Observation f6e3de66-b89d-45fb-87f8-5ed1d36b1be9 · outbound

This paper cites Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.409376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.409376Z digest=sha256:a226f0b72e2a3d56d2222b7e24a2c0daafef98269de0c6b8bbb8ab75c9bb9542

Observation e88a03ae-e068-4566-afa0-7d8f94d08bff · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.532834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.532834Z digest=sha256:44a8031a75d8ad578ca42f556d22bd8c40a4eaf2c242408df5f433e59555516a

Observation fb9bf395-747d-4440-9fa5-1021e695b54d · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards A Rigorous Science of Interpretable Machine Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.662856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.662856Z digest=sha256:91d7e3664b92b026b33be9e5be31a879a36741f6d6ab020025af38afb8401736

Observation cad4fade-f35b-4ddc-a36e-bd417372e751 · outbound

This paper cites MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.729172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.729172Z digest=sha256:1da55843c48d83856f60bccdfe8c064cf44721ab62a6a7a0cfc56c7a014547ff

Observation 682d30aa-73aa-499d-ba2c-27781aec2e4f · outbound

This paper cites The Llama 3 Herd of Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.823979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.823979Z digest=sha256:59399d1bd1a1cadb6d5f7975e38f0f0ec3ced283eceaa29bf6b35a86a3c717b0

Observation 476651c5-f936-451d-a94b-c1b54da0d379 · outbound

This paper cites Alignment faking in large language models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Alignment faking in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.897092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.897092Z digest=sha256:671b69b52595a69a23d2ec4712f15ddfece0cbace54ff4493af4426536e0216c

Observation d54daab9-5337-4536-bacc-7e43417a0aa8 · outbound

This paper cites Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.060892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.060892Z digest=sha256:fc3c6fc6ff68bb762489e055e026a74b56a9a1d7130affbf101e415ea2be47fa

Observation d1ba6d7d-cae9-410d-b8a7-40ee24609032 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Massive Multitask Language Understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.831058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:53.178068Z digest=sha256:3ea61caf8d8994a9e8c1266a60380f1f642a9883657305db7e32d5cb2a2e05b9

Observation 34491400-e581-485d-9654-5ade983b4683 · outbound

This paper cites Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.317192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.317192Z digest=sha256:f2d1127204f3d5d7c393ce8959290137b79ed2075d0fd0f4a7b53c639fcf25d2

Observation e1f3071f-7371-4d66-b1ea-1d360ba676e1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Adam: A Method for Stochastic Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.461451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.461451Z digest=sha256:df0da9ea53b72b146c8cda7fccb261085c17206d83b1e2618d0c5d0abf7bddc4

Observation dfcd6065-a185-47b4-a818-df84799aadc4 · outbound

This paper cites Prover-Verifier Games improve legibility of LLM outputs.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Prover-Verifier Games improve legibility of LLM outputs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.640168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.640168Z digest=sha256:e8700fe7f921852719768dbd6085f69a0116c705969a1543fa264104c712210f

Observation 910bb57f-c351-4715-8196-5f40abf2c201 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.828913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.828913Z digest=sha256:4aa19d68fdd33859416722b36bbce25db81de581c405a707e95bdeebcfb53230

Observation f87c9427-f11e-48c6-b9c4-899d25c0878f · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:04:00.660472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:53.978883Z digest=sha256:80d9bcc1d06c765cf7dfa6e148a646937c0802b75b78ba64950b33dc5a7eed8f

Observation 81be3c40-872a-4346-a988-6ed51f211c33 · outbound

This paper cites Faithful chain-of-thought reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Faithful chain-of-thought reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.394410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.086785Z digest=sha256:1081247ccf8ea6d231072df7ef6c2a81a60877b12c2f578d094757bff115e9bd

Observation f3524604-2834-40fb-af44-93a9b9fd599f · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Frontier Models are Capable of In-context Scheming

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.197470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.197470Z digest=sha256:2901d6ca0658a3daa0de420f96a7c92c3218299f1aa56f57b9ae967cb6beb220

Observation dbef506a-d2e8-4f1b-8ca0-42d10430c43e · outbound

This paper cites Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.287418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.287418Z digest=sha256:f3bbe7176a8dd08b74afe9f82b3a67bc9559c3a4b07882931a44c7c6f0b685ef

Observation 67be5f1f-89f9-49de-b9de-97969602ee1a · outbound

This paper cites Preventing Language Models From Hiding Their Reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.383501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.383501Z digest=sha256:9e5ce7ed4fae53188ece4ff6958d8e9ef0da343e06d5fcd4a948711978461216

Observation 9ffb1065-8644-47bb-8c77-fba73dfbe0ea · outbound

This paper cites J., and Radmard, P.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning J., and Radmard, P

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.213345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.445456Z digest=sha256:6fb9fd7e35de56704fef2af62ce7619975684925c978d25749b8e1ceb21db191

Observation 191524fa-188c-4c1a-9e66-bdf657dc5058 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36: 74952–74965, 2023.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36: 74952–74965, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.025428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.498071Z digest=sha256:bf301b123212ec03aa2ff4cb4c211ef1613b4739f8fca837f26ccc0c2a3c5722

Observation 365824e4-ee44-45c8-b5b5-2e03d6195ba9 · outbound

This paper cites V ., and Zhou, D.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning V ., and Zhou, D

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:59.849231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.613723Z digest=sha256:88455b71d0f098cd12bc6e28425489866d358d0726b8b6b60fa317924d0b927d

Observation 77a23d7b-eb7a-4ae6-a728-63ce0e16f3ee · outbound

This paper cites Stanford professor.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Stanford professor

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:59.485553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.726523Z digest=sha256:c0132f9b2ca9585536e7c50523647c025c10bafbff2cf6bed636e593483f59fb

Observation c5a228f6-007e-4f2a-8cda-5db18b5855ff · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:03:56.417600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.841726Z digest=sha256:0c6f175c53933e16ab41e078d0f97a216b98f6f82112a497b6408e3ed4656a4e

Observation 44717e3c-8013-4a51-99d8-93bce8d1dc22 · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:03:56.224533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:54.998600Z digest=sha256:3a375199c749b51326d79c93a8f4fb974089359f7ee923e8c0c8f7c9f502bdbc

Observation 3373f97b-cddf-4b46-851e-3e81004aa31b · outbound

This paper cites In some cases the bias will be toward the correct answer so in some cases briefly con si de r if the biased answer seems p l a u s i b l e.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning In some cases the bias will be toward the correct answer so in some cases briefly con si de r if the biased answer seems p l a u s i b l e

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:56.014356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:55.131494Z digest=sha256:9cbf9f6be628581b0c113d914f3e34306b10ef4ada81854b0b6d7103d1b30bfa

Observation b6f3221a-d240-4ea6-a402-1de50dfb289e · outbound

This paper cites emotivism.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning emotivism

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:55.825385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:03:55.246559Z digest=sha256:e0c3a4af66bc004e4e04831d84192fb628a48179d98e85101000615d3bed2d49

Observation 8f19acc5-3cfa-4e71-be6a-08e0adfc46c6 · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.437235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.437235Z digest=sha256:63a0bd7a1ddcdd4356a5c5c5d9399a39039f4aa09d81f8a3a6f2867062d6cc29

Pith citing papers

Observation 37af238e-3c66-48b7-b8d8-4a01d16a4f31 · inbound

AI Must not be Fully Autonomous cites this paper.

AI Must not be Fully Autonomous Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:00.347387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:00.347387Z digest=sha256:84cfa1dd5b736c7d0ab77ffca97518d7542a204859397855c21374ba9f36d115

Observation d02bf6ef-1025-4236-a04c-8720533ff381 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.296337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0cf582a78323ea4ab563fd623cdeb848ff75f2c3ba8ee116dc342739ede6c336

Observation 3b8e5177-58e3-46eb-a943-98ca9ca8c001 · inbound

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use cites this paper.

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.602626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T15:38:41.821264Z digest=sha256:66cf50d194a4b1a0f830909c42835f01c46efb07844ecd14bbfbaa2f5829563c

Observation 86df5f34-fae7-474a-989e-fb04c900df8c · inbound

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning cites this paper.

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.985017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:29:34.096277Z digest=sha256:dec1e41dbf158d808c44dec5a55168872489ccb7c09dbe16052090caecda185c

Observation 57d7f00b-157e-4cc2-a731-ce0ac2d98dd7 · inbound

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models cites this paper.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.625126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:81fd21f28426d351d16a07aad5d216be4898ab7b983b1349b928495a5f0ac4b8

Observation fc5cb71a-18b1-49eb-a2ec-5d8a375928a1 · inbound

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings cites this paper.

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:52.207562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:52.207562Z digest=sha256:13e1f7cb96b245967e888271edc358f0dc77b03db3c9ac7072096ff6ca95bbce

Observation 8daac443-4250-474f-8592-87fb3f5eef50 · inbound

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings cites this paper.

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:52.262506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:52.262506Z digest=sha256:61e67445f6460ec5d0142efd53dec82d1a87dcbb7a3cfed35b608142257d6b7b