Pith. sign in

Paper Citation Record · LEDGER

Feedback Loops With Language Models Drive In-Context Reward Hacking

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2402.06627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.06627 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:57:23.024600Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T18:44:28.860347Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d582b9b6-d5ed-4597-91f2-37a359ee5adf · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.862606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:4369222c352ae07a60a6ba7d2fc6c3f71331b51c037323dc4d625cb0bdde65da

Observation bccdd9df-c831-448e-9f5d-c3cdee503a35 · inbound

Thinking beyond the anthropomorphic paradigm benefits LLM research cites this paper.

Thinking beyond the anthropomorphic paradigm benefits LLM research Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T22:21:52.673974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:21:52.673974Z digest=sha256:8a70220e4327af38f221a9f7949479b98464e740b7c990bbf04539d196abbc50

Observation 3e905d17-97ac-4822-9855-bdfcdb3f7df4 · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.461019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:52b131852b13958172f309b8739abd9823fd5e30aeeb3d3ccd48ba6393f25046

Observation fee8b775-2e71-4ce1-b46f-f555e57a8170 · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.928654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:67d1f97f9d53a021f78a5578bcf6d94dd5cbc1bfd8005f844f16f981a596cbfb

Observation 5fc8f8d2-3bc5-4065-acc1-1db253b092fd · inbound

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction cites this paper.

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:49:56.840261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:49:56.840261Z digest=sha256:fbbc7f7d615b99cf3af018f31f32fb3533ca560d26ad61829f9dee55914ab51a

Observation ca149062-f3c4-4f18-b192-c05c9a81c0c0 · inbound

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning cites this paper.

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:57:55.899447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:57:55.899447Z digest=sha256:e0de700e513ead7e83912b752b298c055081e3ac477f448dca152acc2c70c3e5

Observation 636169a7-8160-490b-a12f-060c26b31b6e · inbound

Mitigating LLM biases toward spurious social contexts using direct preference optimization cites this paper.

Mitigating LLM biases toward spurious social contexts using direct preference optimization Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:33:14.744513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:33:04.433907Z digest=sha256:50c626a7cb632848de020bb247432f3d2a91ada63cf9648ad4c98e3fa9634b2e

Observation 207b8bc6-e046-4d22-848b-aea22f51b712 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.675549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:081bbc5c9177652bfde6e31d44f0863f5f3a1f389c3374b80a1e121e1d5b4450

Observation ffc46a41-e1c3-40a8-a160-731b1f3714e3 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.960515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2fd74557bc75d9041ec37f585ef91806bf4b2516434168627ed646e0f9b626df

Observation a3e2b4bd-a5c3-4839-b964-10383a8d73ad · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:08:29.456354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:9223a61495cde97cae7bf789a38d2dc5e7583d121693dddbf8e6dc8142617131

Observation 32f7e0b6-64ed-438b-af0e-fe9c6e80c8c2 · inbound

Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems cites this paper.

Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T08:17:25.689409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:17:25.689409Z digest=sha256:bc1d4534c0c95d0eef1197a6ce20a0d63feb20222a0029337f9415430029758e

Observation 832076a7-7b7e-4022-951a-a04f9ca0be32 · inbound

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs cites this paper.

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T10:57:23.024600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:57:23.024600Z digest=sha256:30fdd96df01f614ed5849eeb58810f797d45f4e08e588017350aff32ea2fa3a5