Pith. sign in

Paper Citation Record · LEDGER

Reward Shaping to Mitigate Reward Hacking in RLHF

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2502.18770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18770 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:41.769204Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d48a086d-209a-4c6b-8926-5cf6b54b797e · inbound

Supervising the search process produces reliable and generalizable information-seeking agents cites this paper.

Supervising the search process produces reliable and generalizable information-seeking agents Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T02:18:27.204122Z digest=sha256:854209a155471957340538725fed3f07d69f811d56802fa0bf7949529eedf6a4

Observation 58c6321b-7665-4423-8b27-86437a460e6d · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:4929f196c9a482b0e486720b2fc4d73fba857f789befc089d035bb3955e69a8f

Observation e947c3dd-4fd3-4e12-876b-9d57f2af9dbf · inbound

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists cites this paper.

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.769204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:41.769204Z digest=sha256:ffefc5fa2404ce508675880ea69166c3fad444f906dc851ee1c01b46082450de

Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.897214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.897214Z digest=sha256:f1f6aa720e40425b30c99f609d7222781305c440c577e7971b0538d2ae0aaec2

Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:84e50759da5f3315b07d5983110046ad33ae08388e699018dd0470e24f7caee9

Observation 1e5c7b63-c8d2-41ef-ba7f-41347e6abf8c · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:33be15588a1228028fbbd7886dc5ada04e72660fbe7c80d69d2ced46d2c015f3

Observation f423b48c-8cd7-4e5b-85fb-08c49e2b8e17 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:6e2bfc02f4138c3fc2cb8a64e28537d9de0e68b26caa70cb6a96bdda27d9fd81

Observation e0aa7ce6-4486-46d5-b0d7-cff2b43fe244 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:6710bb44ce96384b5ed0d82e73c17c467900cf10df716b9dbd9d741d677659f5

Observation a496afa7-0edb-44c0-95fa-ff04bd50f10b · inbound

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning cites this paper.

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:19:41.302164Z digest=sha256:fbf42b927e2f769db9ed6780ea582d852503f97e7fa96cddc5a8686da8e2201d

Observation 31879369-3523-4a1d-82b4-7a711302bdb3 · inbound

Optimal Transport for LLM Reward Modeling from Noisy Preference cites this paper.

Optimal Transport for LLM Reward Modeling from Noisy Preference Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 260

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:b718093ee96dc77bf2b3bcff13a70ee7d275e6b8bd8d6bae2c9dc7195fbb6d8b

Observation 28cb54ed-920c-419c-8d23-2482c8cc6dad · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:89edffb067d4ba15575d04102865e369358d603d4ed5f0b66cd6cc4d515718bc

Observation 4409e985-2190-4e31-b2d7-ef560d706b26 · inbound

Variance-aware Reward Modeling with Anchor Guidance cites this paper.

Variance-aware Reward Modeling with Anchor Guidance Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:11:05.546092Z digest=sha256:393a516c28649274b924efbfef8f5f2b2098a2bede0950f1a6406f9c461b2809

Observation b3fc049c-ece3-453c-8748-63065fc9883b · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:af62dd1ea3a92e9f2181a4f1ebad0425d109d10266fe1cb2ad56d9bf7b409dc3

Observation c98675b8-a2c6-45e8-bd29-ddb4c5b20399 · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:b3f920144da8d1e35555dc618047e78df526390aa4fc81a4073e52be0050805f

Observation f03a1e6c-f8f9-45b8-b88b-552164449677 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:587d874fc7323d1455179717eb6e0a1420a9c487f79baeaee7a8e9be29dcc3a9

Observation c4ec532b-ea35-4495-8772-1c98695e82ae · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:eff02ab29ce96409bc13e4242835f5538e864572bc944ae86313547c14b6f092

Observation 67adf7ce-7aa5-4b17-8c48-c0279bdd6785 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:7dad5e70a70194dbb9023cb5bd877c1ff12d9618d157ac432530af1857b0f1dc

Observation 23180915-ff3b-4ee9-a737-6d9700bc46a9 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:8c0419495d115b414ef991d7417bbb7dda44a85da967f930e0790c110751312c

Observation 542a30f0-8e1a-428d-8eaf-16dd502c6c43 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:c6d30645dc4ff98e28e36eedcd4240d0a0c878d7933476c7806541031f67911c

Observation f25422c1-50d3-4e78-8a25-94556b030c76 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:03cfe0fe3e49e0617d8d4c625925e2ce71ebcefd25bcf92ed80b98db7509ee0b

Observation ae675e5a-4070-45e6-a112-598783ebdaa7 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:aa7f4419298831ae816d1b3ec6c932b05ef7e42061c7a41ee804eee32bc10313

Observation 0c0ba185-f333-4895-adb9-702110f2cef1 · inbound

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning cites this paper.

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:00:37.797967Z digest=sha256:88608aea25ef19d55d041dd0211e5cb0b7be48f8b143764e2f485e48e2c040df