Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2403.03185.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:10:44.092970Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T10:05:40.200096Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 4842cd52-f304-418a-b188-f2215cc35f9f · inbound
CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9fa1578-028d-400b-9587-c93bdb7b4e1e · inbound
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2670194a-a4ac-4f5c-9ff8-66c87f5a78b3 · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f7b982-d8b7-489f-a195-dd984d220137 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 206
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc8c0f6-865f-4505-82a2-67754d8d0c58 · inbound
Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d2ddd8-1325-45e9-89bb-dc941196c020 · inbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 52d9be29-3de1-4273-8d07-116dd9203852 · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 96ff174c-932e-4786-99af-19862dfa8c28 · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c54b6007-68b2-4640-b5a6-73b3efa9d85b · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9340113c-aa56-4a4b-8ee1-294f5aa71c62 · inbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1ab4c228-9c42-4b17-9b41-9cef9a2c692a · inbound
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks? Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e7eaca9b-a368-4663-9a76-c2e67284a7b7 · inbound
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks? Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b1ca5bd7-1be9-443c-9cc7-91332aac482f · inbound
Reframing AGI Confrontation with Off Earth Autonomy Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c85bc5c-0dc3-42ae-98d0-eac8dca4d909 · inbound
MentalThink: Shaping Thoughts in Mental SVG World Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 175
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51def8e8-b57b-4a8b-bbac-6f7c6bdffadd · inbound
Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.