Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.052839Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2502.07599.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.052839Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.454892Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:28:54.031274Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 33145946-9ccf-4395-941c-5093e741f713 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63faaea3-55da-418b-a837-b8b4b46261b3 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Llama 3 model card
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa528fe-52bd-42ff-9cb5-4580d10ffaa2 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Capybara-preferences dataset card
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ba312bda-118d-480d-a798-75fbcb774058 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization A general theoretical paradigm to understand learning from human preferences
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b789d96-2093-4d08-9c54-ec647ad0fa42 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c0063c-2605-4660-85a2-e3cd44243ef6 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ef2061-aab6-4200-b693-268954668416 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 148100a8-e34a-4912-b8ec-1dc0f6cf402f · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Evaluation metrics for language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 48da12f1-a16a-4479-a2c2-b52e41046536 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Deep reinforcement learning from human preferences
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f4ec9a-402a-493b-9787-10a3d55970ce · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Perplexed: Understanding When Large Language Models are Confused
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10a41111-4a58-4256-869c-4013fd5d9ca8 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization UltraFeedback: Boosting language models with high-quality feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bbbc76a2-5801-4144-b6a0-5db54b7d0604 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Enhancing chat language models by scaling high-quality instructional conversations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2396ef46-2d0e-4e60-8dbe-24000f16334e · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c339e4-edd7-42fb-97ce-a3791d515913 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization ORPO: Monolithic Preference Optimization without Reference Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76184030-189c-46ef-918c-4bfd064b0068 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization SDPO: Segment-Level Direct Preference Optimization for Social Agents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b136f837-6d43-4c34-937a-3bc5eb491de4 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa3f7ec-4432-493e-98b6-60b849829c99 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c695b3f-a832-4f63-a452-1d08566350ee · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Entropy Controllable Direct Preference Optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c14b757c-8579-4c70-af67-3c2f211be6be · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f0d12657-a7e6-4a86-b414-7aeefe75b728 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b9d3f8-c4db-4d75-a847-131dd26a31a9 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Pre-dpo: Improving data utilization in direct preference optimization using a guiding reference model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65119daf-0aba-4cb7-a45a-6001e6a64167 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Iterative Reasoning Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc0fbaa-b4f8-4303-a375-4b58990c6ebf · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a6b671e-0090-4735-9254-5d340e7d76f7 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdb08b6-9900-4f8e-a1a3-43257109423c · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Manning, and Chelsea Finn
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63e46b8-c36d-462a-b4de-750bbc12eb94 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7acb93ba-52eb-43f3-aea2-2b3d17db54a2 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Learning to summarize with human feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f9010d-7322-4a94-9531-3db1be88ae8c · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1708aeca-820c-4d29-93d4-736cf1eb84b1 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13350852-ce03-4237-80ce-28b34add8a2a · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7448a94a-bc89-4a02-a58b-c3832af6a8ef · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Rush, and Thomas Wolf
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a88f863-0b9b-4fe8-9650-f4f08dd9e28e · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5983e2f6-f1c3-49aa-999c-9598cfd3d96f · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 631602a3-3634-47d5-a68d-f8891261281f · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization A systematic evaluation of large language models of code
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6b21bd03-5d6a-434e-8f60-6a9969484189 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3486637-e626-491e-b325-76d1a306fcd6 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6cb968-5e4e-46a8-b591-1bf13cb00973 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b26182b0-e13d-47b2-8538-0bb5063d22d7 · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b91271ad-6335-44da-a823-efee30f9354a · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4e3f4637-e232-4ed4-b1e0-d0973781cebb · outbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d220a40-9759-4e58-a761-1d70cbc63392 · inbound
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment DPO-Shift: Shifting the Distribution of Direct Preference Optimization
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec5bba3-48a6-487e-b041-1bc3367e23d9 · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs DPO-Shift: Shifting the Distribution of Direct Preference Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.