Pith. sign in

Paper Citation Record · LEDGER

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

As of 16 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2412.10616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10616 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:55:20.449861Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9be54856-dcf2-4041-92b5-06150ca4a498 · outbound

This paper cites Apply Azuma Hoeffding with its sample average overt + γ terms.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Apply Azuma Hoeffding with its sample average overt + γ terms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:55:20.781548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:55:20.438177Z digest=sha256:5e0d6f117be4952ce640fec3ba400f0a03dc3f6036033280db3b2425dbc47fd6

Observation 45a7588e-81aa-4a12-8963-8613679a9eff · outbound

This paper cites For completeness, we first describe the preference model below: BTL-based Pairwise Preference (Dueling) Model:Consider a decision spaceD.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration For completeness, we first describe the preference model below: BTL-based Pairwise Preference (Dueling) Model:Consider a decision spaceD

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:55:20.771727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:55:20.442345Z digest=sha256:0e43b4987d5e68c956bae8e564e4b43827d8e4a9a682c29375071844f6093b6d

Observation 997853da-6905-415f-8ce0-a13a3c224fb1 · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Dataset Reset Policy Optimization for RLHF

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.344726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.344726Z digest=sha256:c95db6b632534941a13233bf0e48838f21fc6424bc711a54f9b28ccfef9101b7

Observation 4952c080-e234-4766-866f-e4d7774a6aab · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration KTO: Model Alignment as Prospect Theoretic Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.353607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.353607Z digest=sha256:5476d3baa10f89e7fcaa441bd95f291b10eded4bd3d869bd38af18a3ed177e64

Observation 3d883342-ed52-4d67-8bd6-d059d691205a · outbound

This paper cites Foundations of Reinforcement Learning and Interactive Decision Making.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Foundations of Reinforcement Learning and Interactive Decision Making

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.358418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.358418Z digest=sha256:956d8240ee3a1d6015840b12cc903c257457341acb0afead00e2f9ab2e11fa34

Observation d906bcd4-5162-4f85-a9e9-caef7db45874 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Direct Language Model Alignment from Online AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.366546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.366546Z digest=sha256:9b2db095ab00f2744021961bf0ee6d566a034e5d6f972741a1c7c03fb16478bb

Observation f16a6eec-9670-4ec8-95b9-68c29bf7a740 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.375425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.375425Z digest=sha256:8832bf8e7edc6806543bfa6d191c05bef2e5f135da5003a8f875b61b2ccb58af

Observation 76aac503-48f1-477c-b8d0-d529221e938e · outbound

This paper cites Nash Learning from Human Feedback.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Nash Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.379126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.379126Z digest=sha256:dec04be2b15618e4f242ec6766a4bc0dd15ca555af92b0e345bd1e0f383ae98c

Observation 6587a5a8-a90c-488f-9a01-024d1a61ee8c · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.386884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.386884Z digest=sha256:7533624776bd68f0025df129f2b89f90755ef6a8d26f2eb8730e33db475515b9

Observation 063a77eb-5e87-4049-8d7b-eb8c0558ba4b · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.390981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.390981Z digest=sha256:b3baa204f2690383b4cffdf4ab9d90027274a1b7e110944db2e574e06b7aaf72

Observation 0d8f4596-48c9-487e-98d3-a8d74250d44d · outbound

This paper cites The Importance of Online Data: Understanding Preference Fine-tuning via Coverage.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration The Importance of Online Data: Understanding Preference Fine-tuning via Coverage

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.398995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.398995Z digest=sha256:bfecb7c48c960c951efddf702fe45ba4d579f6f9cb037652f5675d7a368c14f9

Observation 8e75d090-37fc-4dae-b31b-168ccefad6f8 · outbound

This paper cites Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:55:20.567068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:55:20.402857Z digest=sha256:f5f2092c4fbbbd537cedd0e4a9b9f87a8ee2724a4c666014712fb254d5bf7be9

Observation f3d36ecf-ffc9-459b-9912-f7762ac9c58b · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Is RLHF More Difficult than Standard RL?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.406641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.406641Z digest=sha256:3d5ddf8fc6a7caf72125730243c41f34d07b250c134b396dafbb9224cb1ea908

Observation ad795cf1-619e-4876-ba73-0065c3d40990 · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Making RL with Preference-based Feedback Efficient via Randomization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.409826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.409826Z digest=sha256:7fa3b68ee2cf6d91a72e62f9e9bc6064866fe2276a9b776b4da87d051d9ff6ac

Observation 99ac4f1f-8235-43d8-9c22-2d2365db660b · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.412841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.412841Z digest=sha256:6350f636b0f121e464dc4761036d25129abb05a64b3b9b7101af016d186e9c89

Observation 99af4eb9-24cf-440a-b6ac-f0f5bb14a70d · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.415875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.415875Z digest=sha256:9fdebc335e7cae4041c0cadeea88d66615d01f0e49fe048d18272bc0c13f0c79

Observation f56f94a3-d0d3-4a03-a3a8-816047500123 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.419053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.419053Z digest=sha256:5b8d9ff8e7c46519f37ab4187bbee0e63da1425fd0ada174de1a3558b06c8588

Observation 162e90e5-149f-4d8d-aa28-12054b696c37 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.422799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.422799Z digest=sha256:72c556a2615623b57b754392f3dc12809f32fa3d418f7dccd27389890466ce4d

Observation 14a9f5f7-209c-4749-896e-6db03eacc56c · outbound

This paper cites GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.426332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.426332Z digest=sha256:200c725892312bc240d74be6ef2c9a5f39f9417b29b4a9da393516a059d03931

Observation 88e81838-4d92-4525-9c47-3872bc995401 · outbound

This paper cites Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.430135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.430135Z digest=sha256:66122d960e97d92a3ec15dde93c035094c3d19769d2f53190109a56fe41154d5

Observation 32e87d64-117b-4fec-88c9-6ca08592d1df · outbound

This paper cites However due to the hybrid nature of our algorithm, we are able to derive concentration lemmas with a faster convergence rate.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration However due to the hybrid nature of our algorithm, we are able to derive concentration lemmas with a faster convergence rate

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:55:20.790917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:55:20.434238Z digest=sha256:f7fe3c8c0fca4efc4ae78a1147f40bbb8d9479e7097578b5496fb2f59af21655

Observation 5c1c55aa-9177-48aa-8dbd-32857511b00c · outbound

This paper cites We will use the abbreviation linB for this feedback model.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration We will use the abbreviation linB for this feedback model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:55:20.760787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:55:20.446202Z digest=sha256:6498e01bef9f1428db4d76bcadabd01532c4b203385cca5246ad9a5c3c4653e4

Observation 845c7a81-4884-4e38-99ce-ab8bb0da3083 · outbound

This paper cites Interested readers are encouraged to go over the proof of Lemma 9 of Saha (2021) to see the proof of lemma 4 above.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Interested readers are encouraged to go over the proof of Lemma 9 of Saha (2021) to see the proof of lemma 4 above

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:55:20.749858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:55:20.449861Z digest=sha256:1568e319d9b444a0c575176ca83cb1f1fda2b4298176cb0185cfb89f05b6f3f9

Observation 5dd43ec3-19c7-4bde-8508-5212ae2fb123 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.340571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.340571Z digest=sha256:08e9c6e1b6b0bfbb257b25da6184fa1bcac47199b2f7f9d759810e1db0cb2882

Observation b6545e01-b8f3-42fa-a2d0-b9a8683d8c68 · outbound

This paper cites Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.382821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.382821Z digest=sha256:0e6c59c5662fb293cd6d95af819a72ab524a6a041baa9ad744e2cb804dbafaa1

Observation c6ffbf91-c158-4f97-8d6d-24765c3f4be9 · outbound

This paper cites Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.349314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.349314Z digest=sha256:0b2dd2019856e9e971313f88c65f938a2f5edebe3d657d6dd8a957b3da4a6187

Observation 5ef03d1e-af95-4195-b97e-9944ce9e8842 · outbound

This paper cites Harnessing Density Ratios for Online Reinforcement Learning.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Harnessing Density Ratios for Online Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.333579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.333579Z digest=sha256:127ae95e39dd3dbf4d36b45c11c5d378acc3a29ebcb5244bd02420eac08050df

Observation c47d63c8-2b0c-490d-9c4d-c6be1bf33edb · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.394661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.394661Z digest=sha256:a7bb720a0b468b65d2b03f6aed8c8e96e45f3f119856f479cf14fbe921bf57fc

Observation 648433e8-5327-4257-bc40-6a092fb608d9 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Statistical Rejection Sampling Improves Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.371083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.371083Z digest=sha256:c04916053adad31a0c8bc801a3fd1eefb2cf02084abd1517d0ea11f12110474b

Observation 557f6430-e525-4608-ac5b-88bd1e4649fd · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.362579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.362579Z digest=sha256:a09ba4da06ae1501f53b0bb9b3dc391a9c7935efde924f19e66ffc83ac2f1595

Observation b4eef0a0-1051-4805-ad27-5df50376f04b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.337568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.337568Z digest=sha256:c8cb8021ad4ba7ee4b38b0b309402d5126027ecbcc6d3084cbaa587bcbab49e9

Pith citing papers

No inbound Pith citation observations are available.