Pith. sign in

Paper Citation Record · LEDGER

Reward Shaping to Mitigate Reward Hacking in RLHF

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2502.18770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18770 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:53:14.002043Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d48a086d-209a-4c6b-8926-5cf6b54b797e · inbound

Supervising the search process produces reliable and generalizable information-seeking agents cites this paper.

Supervising the search process produces reliable and generalizable information-seeking agents Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T02:18:27.204122Z digest=sha256:e171eb810f5df98b98554acf6fbb932ce574da54a0e56ca66b99630a063607ab

Observation 58c6321b-7665-4423-8b27-86437a460e6d · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:c67821d314df543f93208507a6705b5df7b56160e2cb16520afc9d2068d70f67

Observation e947c3dd-4fd3-4e12-876b-9d57f2af9dbf · inbound

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists cites this paper.

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.769204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:41.769204Z digest=sha256:0637dfbbb72b8674cfbb8c7e7f322699038466f0a36f6931df49b5300e735de1

Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.897214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.897214Z digest=sha256:8dbfee4972d913e1f2cba22428890df71cc4e096b678e56dd1f553aa01209d0c

Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:de9cda841a77923183080f986d5fa4cccbd2c39785bcc9e3b8d197b5ff9aa2a6

Observation 1e5c7b63-c8d2-41ef-ba7f-41347e6abf8c · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:3df6f4ef43cff4b5faa93b860582b701971ee5aa59e7fd6192985ebf5835b03a

Observation f423b48c-8cd7-4e5b-85fb-08c49e2b8e17 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:3ee9ab64089c078b41919679858ce2da8067add971a2ce58daa9f4e25365df83

Observation e0aa7ce6-4486-46d5-b0d7-cff2b43fe244 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:bc0e311490992bead6a41fdfd8ffd0cbc905e7c3997ee840d34b675e663c0e9c

Observation a496afa7-0edb-44c0-95fa-ff04bd50f10b · inbound

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning cites this paper.

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T18:19:41.302164Z digest=sha256:9c57ac64b00b8b6697d18fe0e8ed29798c07038fad1437e78bd1527e552777c5

Observation 31879369-3523-4a1d-82b4-7a711302bdb3 · inbound

Optimal Transport for LLM Reward Modeling from Noisy Preference cites this paper.

Optimal Transport for LLM Reward Modeling from Noisy Preference Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 260

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:13f8168af72a83a15a3e534de2f2b1ca9308c4f8fd1296536de8bfb28686474f

Observation 28cb54ed-920c-419c-8d23-2482c8cc6dad · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:d1ee24747baa823270b9787c3fcfdea3e267ac490e4dc08109e5407cc66c640e

Observation 4409e985-2190-4e31-b2d7-ef560d706b26 · inbound

Variance-aware Reward Modeling with Anchor Guidance cites this paper.

Variance-aware Reward Modeling with Anchor Guidance Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T05:11:05.546092Z digest=sha256:0128328577ab6c2083102d7a6c7b2afaeeb4cf91637049793c8ca7cc57d554e1

Observation b3fc049c-ece3-453c-8748-63065fc9883b · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:89bc0435fefe660edf0a0d0fc4be87b1e945e33aeccaa2e70eb6091104a6a4df

Observation c98675b8-a2c6-45e8-bd29-ddb4c5b20399 · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:104f5fd3950cc52f1d16a18a624d41be1c5029f8a03ee5c1c8efb29cd3bcc89a

Observation f03a1e6c-f8f9-45b8-b88b-552164449677 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:729e3d8216d419f7ff0cecf18a7514ec1d31cf6c31b7ce21330a67c51a1a9533

Observation c4ec532b-ea35-4495-8772-1c98695e82ae · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:4eabab2393a8190fbdefe4c8799f069f6a0cbd448422f60a2995206a406661af

Observation 67adf7ce-7aa5-4b17-8c48-c0279bdd6785 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:fe8f3750bcff70b049590a48c49fe7cc83c4a954c84912d860f30b2baf36f054

Observation 23180915-ff3b-4ee9-a737-6d9700bc46a9 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:f9b3b32a18c01904f57d2eb30630cf39552f556d447dc7cac08f563ab386eff2

Observation 542a30f0-8e1a-428d-8eaf-16dd502c6c43 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:269bf8da62ec8057e1c4d1ca7cdab34834f402998a6367ddfa758577faeebc1f

Observation f25422c1-50d3-4e78-8a25-94556b030c76 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:27f5fa2b5ea2fd4693a2c707588caac663934a220d5bef2c8d902aa6640430c6

Observation ae675e5a-4070-45e6-a112-598783ebdaa7 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:eb5de0cd9afe2f0bd9b5be0f41655083baa09fa9684d26454c98a031ae1e3352

Observation 0c0ba185-f333-4895-adb9-702110f2cef1 · inbound

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning cites this paper.

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T21:00:37.797967Z digest=sha256:f715d81f6004b33e6de576a8410d621769d9bf7203e31f3d768cf24783cd7aef

Observation 9056dd3f-3141-49d4-8937-2bf9ac597373 · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.002043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.002043Z digest=sha256:ddb165535efe7bd6edd8d60da792c1f72e6eed02d6ff6f55e58d7553a7a213f9