Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2401.00243.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:26:59.451936Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 048bf3ff-78c4-437b-a3d6-cf4f6004b714 · inbound
Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a0cf1f58-5a3d-435b-90e5-8ed25db26781 · inbound
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8944f8-bf25-48c0-b630-c8a92c4c1f37 · inbound
On the Robustness of Reward Models for Language Model Alignment Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07acfe67-c226-46b8-9be1-2d768943a270 · inbound
Towards Reliable, Uncertainty-Aware Alignment Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed2f1f9-26f6-4894-b619-9d57d5166234 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d70d49c-04c2-4d1a-9053-0af56be2a8a1 · inbound
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 49710199-27b5-495b-918b-c2b02f854056 · inbound
A Unifying Lens on Reward Uncertainty in RLHF Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 71b7d724-a887-4a74-8164-0c2720d079cb · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dc9bb22a-5858-4d1a-b6c9-656377e7607d · inbound
Reinforcement learning to improve large language model-based automated code compliance systems Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18f85772-110c-474e-a952-ca123ce32896 · inbound
Internal Pluralism and the Limits of Pairwise Comparisons Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 179
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bfe821-9a1e-4995-bf8a-a21bd18b574e · inbound
Internal Pluralism and the Limits of Pairwise Comparisons Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Reference 276
Source-reported events for the cited work
Unavailable: canonical work link unavailable.