Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:27:11.573917Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2502.04567.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:27:11.573917Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:58.461665Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T05:36:40.511807Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 09ff3f3f-553d-4098-9317-addc4c038319 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator G., Guo, Z
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f98236-acca-44d5-92ad-5d844804e912 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator E., and Nocedal, J
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5467d72f-8252-4146-b306-2a0d4123fe8a · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4898ca2-5e47-4c0f-a946-6fc3ac606d25 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Noise contrastive alignment of language models with explicit rewards
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7346bb9c-6aa5-488e-b262-cc6f8f0b2e1b · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Towards Improved Preference Optimization Pipeline: from Data Generation to Budget-Controlled Regularization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 75cdeb0f-4ad1-4c10-984a-685cf5c4f2f1 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Ultrafeedback: Boosting language models with high-quality feedback, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0d6ab18-1fa7-465d-b33f-6b2767e23a04 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Anchored preference optimization and contrastive revisions: Addressing underspecification in alignment
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39445c87-3d7c-40dc-b8e9-7e36c312a16c · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc34f322-dd1d-470a-bd37-86f452694010 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88f3c83e-1990-48cd-94d2-8d307d1c65c6 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator KTO: Model Alignment as Prospect Theoretic Optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dcb33f1-8483-413f-b250-8e4b0ab79b67 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Aligning language models with preferences through f-divergence minimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 35a8d537-a80e-4db4-a8d6-ec6e39ef4f5e · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc708a27-0e70-4be8-a02d-77803f193388 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator and Hyv \"a rinen, A
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e29e08-df1f-4a04-8eb8-4cd62e0c91e4 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f27505-97de-4b92-8df5-5e535e5f01c2 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Towards Efficient Exact Optimization of Language Model Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d54cf7-f402-4d13-b4fd-a5817cdb209c · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Binary Classifier Optimization for Large Language Model Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe46b0e-7036-44ba-9b12-059c391530f2 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator H., Gonzalez, J
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b3714b-1694-452f-b7c2-092cb8ee9d0a · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a7362e-cc12-45d8-a1cf-48eb939c46a1 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator E., and Stoica, I
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation df2d7f18-845b-4a78-b1e6-f792ebd39473 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad8a705-5d0e-4ce1-bad6-2c9cfea9a834 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88278b66-2fb6-4044-b725-e202895d8152 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f9d9fca-3083-4181-b1cc-b7ffc7a440d3 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Elements of Sequential Monte Carlo
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44bc280-46d5-4d8a-a633-4e725dfccac6 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator On the connection between noise-contrastive estimation and contrastive divergence
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cc703708-848f-4776-9c6b-79de8ddd461e · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Training language models to follow instructions with human feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad16160-6c38-41bd-bbcb-c7b5c618f451 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667a8f3b-64e8-4391-a4af-d0f26ba1f572 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator and Schaal, S
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14537d1c-3b73-45bc-a6d9-917caa9cc877 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d446647-df40-4f84-ba8e-3d50dc7770ad · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator D., Ermon, S., and Finn, C
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95964de6-65b9-4c72-a694-e5fa675a7a98 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e346559-95c2-4063-973b-b7676843897e · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Proximal Policy Optimization Algorithms
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc894399-ce8f-495a-af70-8e1f93fa99de · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b788fd2a-d719-4b25-9826-3906610d7ff4 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Understanding the performance gap between online and offline alignment algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94169a91-f4f2-4441-9332-0d867585e7bf · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Zephyr: Direct Distillation of LM Alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a80611-c1ba-4d9f-8571-f9768e024683 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Self-Play Preference Optimization for Language Model Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23606bc7-2505-4ebd-8220-6702321e2409 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88fccbc1-22cf-488e-b8fb-5bc9d8170d15 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9d8cecb-6904-4a76-be1a-5d3f8f3a8008 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Fine-Tuning Language Models from Human Preferences
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5528dc08-a4d6-48f8-9b59-19331eba3f92 · outbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator write newline
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8216cd18-29cc-4630-88ef-433fb7559ef1 · inbound
York's Cavity Formalism and Quantum Modified Thermodynamics of (2+1)D Black Holes Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268757a8-56ca-4ae9-bb86-dfe3e9422386 · inbound
DeepEyesV2: Toward Agentic Multimodal Model Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a4553154-f928-474f-9534-aa324d786347 · inbound
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.