Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2502.18770.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:41.769204Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation d48a086d-209a-4c6b-8926-5cf6b54b797e · inbound
Supervising the search process produces reliable and generalizable information-seeking agents Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58c6321b-7665-4423-8b27-86437a460e6d · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 170
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e947c3dd-4fd3-4e12-876b-9d57f2af9dbf · inbound
Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · inbound
RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · inbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e5c7b63-c8d2-41ef-ba7f-41347e6abf8c · inbound
Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f423b48c-8cd7-4e5b-85fb-08c49e2b8e17 · inbound
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0aa7ce6-4486-46d5-b0d7-cff2b43fe244 · inbound
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a496afa7-0edb-44c0-95fa-ff04bd50f10b · inbound
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31879369-3523-4a1d-82b4-7a711302bdb3 · inbound
Optimal Transport for LLM Reward Modeling from Noisy Preference Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 260
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28cb54ed-920c-419c-8d23-2482c8cc6dad · inbound
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4409e985-2190-4e31-b2d7-ef560d706b26 · inbound
Variance-aware Reward Modeling with Anchor Guidance Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3fc049c-ece3-453c-8748-63065fc9883b · inbound
Reward Hacking in Rubric-Based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c98675b8-a2c6-45e8-bd29-ddb4c5b20399 · inbound
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f03a1e6c-f8f9-45b8-b88b-552164449677 · inbound
General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4ec532b-ea35-4495-8772-1c98695e82ae · inbound
General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67adf7ce-7aa5-4b17-8c48-c0279bdd6785 · inbound
General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 23180915-ff3b-4ee9-a737-6d9700bc46a9 · inbound
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 542a30f0-8e1a-428d-8eaf-16dd502c6c43 · inbound
Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f25422c1-50d3-4e78-8a25-94556b030c76 · inbound
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae675e5a-4070-45e6-a112-598783ebdaa7 · inbound
Uncertainty-Aware Reward Modeling for Stable RLHF Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c0ba185-f333-4895-adb9-702110f2cef1 · inbound
OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.