Pith. sign in

Paper Citation Record · LEDGER

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 7 inbound Pith citation observations for arXiv:2506.22777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22777 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:03:55.246559Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:51:52.262506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:34:47.983586Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f21d7222-a4fa-4ed0-83a0-10f82d45ea24 · outbound

This paper cites Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:50.441783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:50.441783Z digest=sha256:99c463d241ccc91989f4a1856019d984bb109caf9a942997b87864e45da68f52

Observation ff74f8e3-c968-4801-8c60-1943027adc10 · outbound

This paper cites Claude 3.7 sonnet system card, 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Claude 3.7 sonnet system card, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.747302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:50.497422Z digest=sha256:895582eeb26cd5b52d4c8cb493ced9ccfb2b6039aabcbb56359ba4b13598fcf4

Observation 302fb93a-fc55-4419-bffc-fe5132b7ef40 · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning System card: Claude opus 4 & claude sonnet 4, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.576798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:50.661471Z digest=sha256:79689a79d04ab8a157e5575613333899050f6c8a01cf2feb98fab96b341f8eda

Observation 2ad3bb70-9718-4362-9b7b-5b5f34ad5512 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:50.831268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:50.831268Z digest=sha256:ffeaedb8b93a6bbee24d5955951ec5ee8233dc855efa7f5259a7c1565d2d29d5

Observation 239f0235-1f8c-40fb-b310-915a321296d4 · outbound

This paper cites Do models say what they learn?, 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do models say what they learn?, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.427022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:50.892798Z digest=sha256:2db8900117af4ea2bbab418a76a5583eaa35dcfda4fc8126774f82df87bcc88d

Observation c841f2e4-0213-4768-924c-436d2805774b · outbound

This paper cites CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring, May 2025.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring, May 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:01.189283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:51.084035Z digest=sha256:d772c426d0ebb83be5d71a2d49e3ee2c6634cbc3d45b3d5fb5e20092eccdbe44

Observation a3b04bbd-4df7-4467-88d2-39c5d647e735 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.211271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.211271Z digest=sha256:d273c80b00afa06ea09afcc749dfdf63adac010471d243b1bffda2eb816e0c07

Observation 616b1434-3fc4-4f52-b899-1ce6c2dff02c · outbound

This paper cites C., Macar, U., Nanda, N., and Conmy, A.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning C., Macar, U., Nanda, N., and Conmy, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.996924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:51.342380Z digest=sha256:14cecf60aef103300f0982163bda8c78374122606db91839eb62a2d2484ca83d

Observation 1bde6da0-f4f3-4a8c-9f28-1afffd67c247 · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Progress on Scalable Oversight for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.583235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.583235Z digest=sha256:0f7865dab3d0c7297c5d56d0cdaee8340682ada3d08d7d199a4f4c8c5236b2b7

Observation 54d38c93-7ef8-4d85-8794-6adf0cf2c45a · outbound

This paper cites Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.721948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.721948Z digest=sha256:a0bfc7b3e8c53970dc9e5e321d973cd255778016cf784ce99aa647d0db1abc08

Observation 34c38ac9-afcd-4cee-aaa6-c721edec54f4 · outbound

This paper cites Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.842152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.842152Z digest=sha256:53384f43d6d2f55f5907e2c467ed3919aa4c44beadf84439418675b251d37a08

Observation 2170fee5-fc74-4209-a415-4a639f0c0b89 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.958119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.958119Z digest=sha256:8904370251ab008273e6d2f283a5039706e932eb3968a55cc0839c81c3ae20fe

Observation dc59f3ac-4435-40e3-8787-ad58d3430371 · outbound

This paper cites Are DeepSeek R1 And Other Reasoning Models More Faithful?.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Are DeepSeek R1 And Other Reasoning Models More Faithful?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.114166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.114166Z digest=sha256:22ee04ae26b1ed7be1ae86882f67766200058b97acdbfb368d9b381edbb38131

Observation 246b518d-1ab1-4735-b78f-3311d85ac3f2 · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.256909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.256909Z digest=sha256:432757aa50b3b425e05af1b1973027f114f3ae89bc6149435b739d6cb17a3df7

Observation f6e3de66-b89d-45fb-87f8-5ed1d36b1be9 · outbound

This paper cites Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.409376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.409376Z digest=sha256:00d2bab91bc83c8915aeb60925671c762b76261393d63e9b49696e8418401d88

Observation e88a03ae-e068-4566-afa0-7d8f94d08bff · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.532834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.532834Z digest=sha256:d3512e7937f82cdf9eec6b5baa2a8f7892d4006b3ba4f3e8ef155b82b0f773af

Observation fb9bf395-747d-4440-9fa5-1021e695b54d · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards A Rigorous Science of Interpretable Machine Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.662856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.662856Z digest=sha256:7f90ca429d8c253b6fe32170e863f38c0a76f5d74383342060985a197961ec1a

Observation cad4fade-f35b-4ddc-a36e-bd417372e751 · outbound

This paper cites MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.729172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.729172Z digest=sha256:6fc00d89f0093fe2e9f970bfddaf0d8ead9b8d7ce0b19b6472bcaeca742eec91

Observation 682d30aa-73aa-499d-ba2c-27781aec2e4f · outbound

This paper cites The Llama 3 Herd of Models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.823979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.823979Z digest=sha256:8f5ea77f11e75d367aacabb82977be6211e36d5f29d8d9d80cc908b46d5f2ac3

Observation 476651c5-f936-451d-a94b-c1b54da0d379 · outbound

This paper cites Alignment faking in large language models.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Alignment faking in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:52.897092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:52.897092Z digest=sha256:df2bc53bdeaf099acfa9fa52822332b6dda03d25bdfcfb1fc26d27f05ff80c70

Observation d54daab9-5337-4536-bacc-7e43417a0aa8 · outbound

This paper cites Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.060892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.060892Z digest=sha256:aefbeb877d680f1014720be2e94f7b58cdd4f264bdab1f203dc30c99e6a7b785

Observation d1ba6d7d-cae9-410d-b8a7-40ee24609032 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Massive Multitask Language Understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.831058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:53.178068Z digest=sha256:33fd5d97fc296e8ae60144fe71b5311b6ddc86732f297148005fa54f02309847

Observation 34491400-e581-485d-9654-5ade983b4683 · outbound

This paper cites Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.317192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.317192Z digest=sha256:fb402cb4867d7619a80302a1569fc5e9a9fe68b189d0bc00bf904cb59aaf9ae1

Observation e1f3071f-7371-4d66-b1ea-1d360ba676e1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Adam: A Method for Stochastic Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.461451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.461451Z digest=sha256:dd51b94c662b47db427ca8f9464ac37b77d368987d483701aef8c925baadd18d

Observation dfcd6065-a185-47b4-a818-df84799aadc4 · outbound

This paper cites Prover-Verifier Games improve legibility of LLM outputs.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Prover-Verifier Games improve legibility of LLM outputs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.640168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.640168Z digest=sha256:5bca8b18dfd560b5a73fb9b14ecf4b6c6d432876c4af0391fb32f1c9f9bf269f

Observation 910bb57f-c351-4715-8196-5f40abf2c201 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:53.828913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:53.828913Z digest=sha256:bfd2115c0c62934220dec6f3e9e826190e15cce5871383962e0b4a0645343634

Observation f87c9427-f11e-48c6-b9c4-899d25c0878f · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:04:00.660472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:53.978883Z digest=sha256:08967ea3fca0954dd0baee9117b4707b1173df2b5c9cdfe0c84a88033fef6c33

Observation 81be3c40-872a-4346-a988-6ed51f211c33 · outbound

This paper cites Faithful chain-of-thought reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Faithful chain-of-thought reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.394410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.086785Z digest=sha256:b43af72ed6a5d588cdd6c63ccfbf7aab3b3b258c4bccda4fab7350f3d469c63f

Observation f3524604-2834-40fb-af44-93a9b9fd599f · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Frontier Models are Capable of In-context Scheming

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.197470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.197470Z digest=sha256:befbe3b49e3045fb627f308358111956354650b5874c5677d55cd684e1f9050c

Observation dbef506a-d2e8-4f1b-8ca0-42d10430c43e · outbound

This paper cites Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.287418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.287418Z digest=sha256:d5e62dbb66303e54a6d252faff58c307fff99daa44581a2f164c4fc0b6af3eec

Observation 67be5f1f-89f9-49de-b9de-97969602ee1a · outbound

This paper cites Preventing Language Models From Hiding Their Reasoning.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.383501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.383501Z digest=sha256:cb8e8cbcd0821a4f9a2d3d2f370ed41bd91650539f7c844bfbb0e0c19223bf3d

Observation 9ffb1065-8644-47bb-8c77-fba73dfbe0ea · outbound

This paper cites J., and Radmard, P.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning J., and Radmard, P

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.213345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.445456Z digest=sha256:3ef55911bc6dc81b7e3cafd8545e52d5217b1f1326fb649bfef8ffcefdbfeb05

Observation 191524fa-188c-4c1a-9e66-bdf657dc5058 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36: 74952–74965, 2023.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36: 74952–74965, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:04:00.025428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.498071Z digest=sha256:70dbcefc33745e0d3dca994f963698278e36f33276421ed367e01db1713c4349

Observation 365824e4-ee44-45c8-b5b5-2e03d6195ba9 · outbound

This paper cites V ., and Zhou, D.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning V ., and Zhou, D

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:59.849231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.613723Z digest=sha256:919f251901e59a3eae65609247c005fb79163082b04ac489622dc2a447fccffe

Observation 77a23d7b-eb7a-4ae6-a728-63ce0e16f3ee · outbound

This paper cites Stanford professor.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Stanford professor

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:59.485553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.726523Z digest=sha256:013a777169d8f22321ad801c80540b1e2c5ea252eda2c18e6a6d8c8371f44159

Observation c5a228f6-007e-4f2a-8cda-5db18b5855ff · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:03:56.417600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.841726Z digest=sha256:6eb17e4a412c5347527b8fd0191a3137100b8723b17d3d5a5fa21547bd4ac956

Observation 44717e3c-8013-4a51-99d8-93bce8d1dc22 · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:03:56.224533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:54.998600Z digest=sha256:db04c6f0422ba49aba205e14dbbf7ec90ae9e689252a741d0699b34f72813ff4

Observation 3373f97b-cddf-4b46-851e-3e81004aa31b · outbound

This paper cites In some cases the bias will be toward the correct answer so in some cases briefly con si de r if the biased answer seems p l a u s i b l e.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning In some cases the bias will be toward the correct answer so in some cases briefly con si de r if the biased answer seems p l a u s i b l e

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:56.014356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:55.131494Z digest=sha256:d011849809db748be454401a48128e93013855605bb5eaab80a6afcd4ed5c24f

Observation b6f3221a-d240-4ea6-a402-1de50dfb289e · outbound

This paper cites emotivism.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning emotivism

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:03:55.825385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:03:55.246559Z digest=sha256:8c22a630ce84041cbf504b9444d3753fccfbe861ea8347bfd44e7b33344fcd69

Observation 8f19acc5-3cfa-4e71-be6a-08e0adfc46c6 · outbound

This paper cites an unresolved cited work.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.437235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.437235Z digest=sha256:2e4a18aa3a516340362fa4f189ffd1096c942dd580067fa77e9d4414a2f42f11

Pith citing papers

Observation 37af238e-3c66-48b7-b8d8-4a01d16a4f31 · inbound

AI Must not be Fully Autonomous cites this paper.

AI Must not be Fully Autonomous Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:00.347387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:00.347387Z digest=sha256:fa0eba11f02cf50e2950527e04500a5a0799e53e9360eaf5c7c7de43f6d02601

Observation d02bf6ef-1025-4236-a04c-8720533ff381 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.296337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c246452351ed4e6c6be4bc5f997e554809da2eae9ca41b8f62c7609912fa7738

Observation 3b8e5177-58e3-46eb-a943-98ca9ca8c001 · inbound

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use cites this paper.

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.602626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T15:38:41.821264Z digest=sha256:e020bf50250c13943e6550c8ccdf5c98e4454ac963452a931ebbb71098eac800

Observation 86df5f34-fae7-474a-989e-fb04c900df8c · inbound

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning cites this paper.

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.985017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T15:29:34.096277Z digest=sha256:a9800403b4e247d91344796266471b82e47e1624105f6ff12729353b20020c76

Observation 57d7f00b-157e-4cc2-a731-ce0ac2d98dd7 · inbound

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models cites this paper.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.625126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:dda45a649f2ddb6ce2e60ae918207676bfcb0e70af3a73e784f162a3ce8a593d

Observation fc5cb71a-18b1-49eb-a2ec-5d8a375928a1 · inbound

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings cites this paper.

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:52.207562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:52.207562Z digest=sha256:23da08d04cd7c567cd69593699d6ccdc65fb2abba0218a2c32615f40da39300b

Observation 8daac443-4250-474f-8592-87fb3f5eef50 · inbound

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings cites this paper.

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:52.262506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:52.262506Z digest=sha256:aa4efe5ea512928db2f1e328b4aba83ac75ae274b1a16c9ebb8e2e16261ade3a