Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2312.11456.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:50.431221Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T08:07:45.294768Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation be79fbfa-d555-494a-8c4a-1fae5e1094ea · inbound
Self-Rewarding Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb628c24-4853-460c-a3d4-056fa4cffffd · inbound
Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 4344
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c925f773-402e-4005-aad0-6c67f4ba00c6 · inbound
Learning a Pessimistic Reward Model in RLHF Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26b3256-2698-4ae6-8ffa-1f03b099f100 · inbound
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 133
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd6201e-da41-4d52-bed5-5e89a29b156c · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d633d1f9-2b01-4aa8-929c-e316d7ca4f71 · inbound
Thompson Sampling in Online RLHF with General Function Approximation Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738674b3-804c-466f-8cb8-2176cf94dc03 · inbound
Aligning Large Language Models with Implicit Preferences from User-Generated Content Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cbcdcb0-37db-4fba-b52b-61e85f8acf78 · inbound
Boosting LLM Reasoning via Spontaneous Self-Correction Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aadb1a06-6ba7-44f1-8ad0-d2b140535341 · inbound
Bridging Offline and Online Reinforcement Learning for LLMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a5f7fda-8506-4f53-83f0-03e83d0c13f5 · inbound
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation caa7a167-a786-40b8-ad11-8e33e6c466dd · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce8968b-b7ec-4828-97f6-67d3d1edf231 · inbound
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9f47699-f286-41a3-83c0-869043ed15e1 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca620d9e-6f09-4fc7-bfbe-5a69903cbd8f · inbound
Outcome-based Exploration for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46296f8-9f42-4f61-bc8a-4a54339c294d · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 207
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e2a028-9ce8-437c-82c9-b871889ae3d1 · inbound
T-TAMER: Provably Taming Trade-offs in ML Serving Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a187e3c4-ed2a-4fd9-bc9c-25ad2c792f07 · inbound
Multiplayer Nash Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29f80c5d-e9d3-4b69-a6ca-ff32d85aa70e · inbound
Improved Bounds for Private and Robust Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f384787-e9db-4a2a-bf85-b581734215fd · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51151045-dabe-41db-85a2-582cede2776b · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11128f59-79a4-4dc7-90d6-57e1d37c1fe3 · inbound
Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40f61a43-74ac-4b71-8338-b529d1583785 · inbound
IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 110138e7-d015-4b0d-a398-f9f63980401c · inbound
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ef6e794-1ad0-4748-a798-b34efe831837 · inbound
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ba08c5f-0a38-41ec-8355-2b6b0715ab24 · inbound
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb3dd61e-1eb5-4ff0-841d-65edf791a5a6 · inbound
HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07b58e96-e25d-472f-80e9-06479c7bbd52 · inbound
The Power of Test-Time Training for Approximate Sampling Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cec4c83d-8ba5-45b8-b569-d375ceb5c8f1 · inbound
When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05b50ba8-3813-405f-8765-7a6df5cff1c2 · inbound
Subjective Risk Decomposition: A New View for Uncertainty Quantification Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab20eedc-4fd4-4669-abb6-94f61bdc20e0 · inbound
Normalized Rewards for Preference Optimization Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.