Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.080296Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2502.06861.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.080296Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T13:06:54.002248Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T13:10:10.404876Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 18182556-cba3-4140-bd73-4b789af87ceb · outbound
Design Considerations in Offline Preference-based RL Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c30b055-fb90-4bf3-8124-fbb6fabce50c · outbound
Design Considerations in Offline Preference-based RL Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 89e4aabb-baae-45b5-af0f-9e2d9e9d6917 · outbound
Design Considerations in Offline Preference-based RL Robust Preference Optimization through Reward Model Distillation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b3bc5c-75f7-4b8e-bf00-8d71f0b107ff · outbound
Design Considerations in Offline Preference-based RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e476658-38ac-4c5e-8c3c-b49441490b4c · outbound
Design Considerations in Offline Preference-based RL Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba1fac7-9632-4932-a5e4-4a15ee9fb486 · outbound
Design Considerations in Offline Preference-based RL Nash Learning from Human Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94c4030-9263-4763-91fc-8508caf05beb · outbound
Design Considerations in Offline Preference-based RL Disentangling Length from Quality in Direct Preference Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f84b809-9c0b-464b-9ae6-7e16dff56c16 · outbound
Design Considerations in Offline Preference-based RL Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4639f06b-92d8-410e-aac7-ed4c6e12b946 · outbound
Design Considerations in Offline Preference-based RL Gemini: A Family of Highly Capable Multimodal Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64d54ee-d14d-4af9-a98e-368242dbd36e · outbound
Design Considerations in Offline Preference-based RL Is RLHF More Difficult than Standard RL?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c00348f-ff9c-439e-a0ce-a04c3d03c991 · outbound
Design Considerations in Offline Preference-based RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5641bb44-ced0-44c8-8857-b9c0b6d82cc5 · outbound
Design Considerations in Offline Preference-based RL The optimizer used is Adafactor with learning rate that is constant with a linear warm-up for 2000 steps and a base rate of 1e −
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02e5e5a0-dffe-4b30-822d-13b599384c68 · outbound
Design Considerations in Offline Preference-based RL Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8739783-b312-4542-905a-f34c8373c5f4 · outbound
Design Considerations in Offline Preference-based RL Calibrating Sequence likelihood Improves Conditional Language Generation
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a37fcd4e-f981-4ab9-8bd2-ce2ed2866e01 · outbound
Design Considerations in Offline Preference-based RL On Regularization via Early Stopping for Least Squares Regression
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caabea31-3e40-41bf-b61e-2ab71ce38f96 · outbound
Design Considerations in Offline Preference-based RL SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d202202d-2a25-428d-938b-5161964e17b7 · outbound
Design Considerations in Offline Preference-based RL KTO: Model Alignment as Prospect Theoretic Optimization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1daba942-17dd-4a26-bfe4-8731dcd449fd · outbound
Design Considerations in Offline Preference-based RL A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1f0652-e606-4e17-8087-744fdbbf5409 · outbound
Design Considerations in Offline Preference-based RL Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d56662-a529-4de4-8ddc-6401e9a9a7d6 · outbound
Design Considerations in Offline Preference-based RL Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c3038c-a61c-4307-9b25-f209c322ef1c · outbound
Design Considerations in Offline Preference-based RL Direct Preference Optimization with an Offset
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe1c548-1e61-48a0-9c5f-c0a9d67d8b9e · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Design Considerations in Offline Preference-based RL
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a5f2044-b3f1-4b49-9a79-dc6d9ad3e799 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Design Considerations in Offline Preference-based RL
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.