Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:55:20.449861Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.10616.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:55:20.449861Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T01:00:59.114689Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T01:00:59.301176Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9be54856-dcf2-4041-92b5-06150ca4a498 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Apply Azuma Hoeffding with its sample average overt + γ terms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 45a7588e-81aa-4a12-8963-8613679a9eff · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration For completeness, we first describe the preference model below: BTL-based Pairwise Preference (Dueling) Model:Consider a decision spaceD
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 997853da-6905-415f-8ce0-a13a3c224fb1 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Dataset Reset Policy Optimization for RLHF
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4952c080-e234-4766-866f-e4d7774a6aab · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration KTO: Model Alignment as Prospect Theoretic Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d883342-ed52-4d67-8bd6-d059d691205a · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Foundations of Reinforcement Learning and Interactive Decision Making
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d906bcd4-5162-4f85-a9e9-caef7db45874 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Direct Language Model Alignment from Online AI Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16a6eec-9670-4ec8-95b9-68c29bf7a740 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76aac503-48f1-477c-b8d0-d529221e938e · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Nash Learning from Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6587a5a8-a90c-488f-9a01-024d1a61ee8c · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 063a77eb-5e87-4049-8d7b-eb8c0558ba4b · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8f4596-48c9-487e-98d3-a8d74250d44d · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e75d090-37fc-4dae-b31b-168ccefad6f8 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f3d36ecf-ffc9-459b-9912-f7762ac9c58b · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Is RLHF More Difficult than Standard RL?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad795cf1-619e-4876-ba73-0065c3d40990 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Making RL with Preference-based Feedback Efficient via Randomization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ac4f1f-8235-43d8-9c22-2d2365db660b · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99af4eb9-24cf-440a-b6ac-f0f5bb14a70d · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56f94a3-d0d3-4a03-a3a8-816047500123 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162e90e5-149f-4d8d-aa28-12054b696c37 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a9f5f7-209c-4749-896e-6db03eacc56c · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88e81838-4d92-4525-9c47-3872bc995401 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e87d64-117b-4fec-88c9-6ca08592d1df · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration However due to the hybrid nature of our algorithm, we are able to derive concentration lemmas with a faster convergence rate
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5c1c55aa-9177-48aa-8dbd-32857511b00c · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration We will use the abbreviation linB for this feedback model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 845c7a81-4884-4e38-99ce-ab8bb0da3083 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Interested readers are encouraged to go over the proof of Lemma 9 of Saha (2021) to see the proof of lemma 4 above
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5dd43ec3-19c7-4bde-8508-5212ae2fb123 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6545e01-b8f3-42fa-a2d0-b9a8683d8c68 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ffbf91-c158-4f97-8d6d-24765c3f4be9 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef03d1e-af95-4195-b97e-9944ce9e8842 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Harnessing Density Ratios for Online Reinforcement Learning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47d63c8-2b0c-490d-9c4d-c6be1bf33edb · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648433e8-5327-4257-bc40-6a092fb608d9 · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Statistical Rejection Sampling Improves Preference Optimization
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 557f6430-e525-4608-ac5b-88bd1e4649fd · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration REBEL: Reinforcement Learning via Regressing Relative Rewards
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4eef0a0-1051-4805-ad27-5df50376f04b · outbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99eb5fc0-a8a6-4d90-b54b-f04130792ae2 · inbound
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.