Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2409.15360.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:30:36.307631Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:37:30.559409Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 213e1a0e-5b42-4517-ad6c-342bee73de6c · inbound
Knowledge Boundary of Large Language Models: A Survey Reward-Robust RLHF in LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127c338a-825f-433d-9939-9d844b6dc90d · inbound
Baichuan4-Finance Technical Report Reward-Robust RLHF in LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15e9c54-7036-46d7-8e33-9a41296b7a8d · inbound
Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach Reward-Robust RLHF in LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1327b6b7-25c7-43f0-9cc9-2622a2a96aa8 · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward-Robust RLHF in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b578d7-c0c7-49b4-ba62-867ceef01ac0 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward-Robust RLHF in LLMs
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a401b329-e1b9-43e1-9961-592a1ba8383b · inbound
Learning to Control Summaries with Score Ranking Reward-Robust RLHF in LLMs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1945aa31-450c-4866-a9e7-61e18991e1c2 · inbound
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Reward-Robust RLHF in LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8491470d-6bcd-448d-9d64-fda26e931bf3 · inbound
PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs Reward-Robust RLHF in LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6e893e5-515f-43ea-b111-8a29f7588a23 · inbound
Efficient Preference Poisoning Attack on Offline RLHF Reward-Robust RLHF in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 02fd9cc7-723c-42cf-815a-bb74bb9a983a · inbound
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR Reward-Robust RLHF in LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b1698c6-85d1-4d09-996a-25ae3ab407e6 · inbound
A Unifying Lens on Reward Uncertainty in RLHF Reward-Robust RLHF in LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fdc39863-b19b-451a-9a85-2277f186b268 · inbound
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning Reward-Robust RLHF in LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c20bfd58-f270-4d15-b8ef-ee2bed78555f · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reward-Robust RLHF in LLMs
Reference 233
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation be211758-972e-4e32-bb4f-70a2031c0cdf · inbound
Multimodal Reward Hacking in Reinforcement Learning Reward-Robust RLHF in LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5acc56-f3af-44d1-b7b0-0daffa70db87 · inbound
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text Reward-Robust RLHF in LLMs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b01754-a1cc-4854-a647-b298e0e23486 · inbound
Robust General Utility for Reinforcement Learning Reward-Robust RLHF in LLMs
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.