Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:56:28.769248Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 3 inbound Pith citation observations for arXiv:2502.00203.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:56:28.769248Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-23T03:50:03.720389Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T03:52:29.440835Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69e45135-5d22-49a7-bf09-00d5fa0e46f6 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3627262-c55a-41e9-b908-abca547f4dd9 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e749a7b-8e54-4a70-b955-848fab332b7b · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment KTO: Model Alignment as Prospect Theoretic Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f408ca-681e-4cd4-8491-8de269854632 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Robust Preference Optimization through Reward Model Distillation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab1b5b4-5906-4666-9f96-4d2cedd5778c · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Measuring Massive Multitask Language Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a1f14d-6478-4baf-afa2-647e70f491c5 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment RewardBench: Evaluating Reward Models for Language Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3f52cc-106b-49cc-a082-07622696c53c · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ddadaa7-79ca-4955-90c9-4269033b9fd9 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Understanding Reference Policies in Direct Preference Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b6e860-e0be-41eb-967f-f367df79c73a · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5fae76-ad93-4da6-9fe2-5c56c71041ef · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c0fcfeb-f141-4e29-a27c-39268687e145 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Nemotron-4 340B Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 384bfe6f-d557-4d0b-9b7d-d8d4b5c0f8af · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment GPT-4 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b9be56-243b-4c62-9fa7-7ebc6fdeee8b · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc797453-adac-4bdb-ab5e-8508b2d32668 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 782e18f3-73c0-4f7e-9288-22e51973d69b · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e702691-ad86-4d0c-9497-413ef50c3867 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12010ffa-1fd9-42e5-addc-0cff30feee9b · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9207d0b3-f7ff-4bed-bf8f-158b6cd46bd3 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Derivation of the distance functions for the multi-response scenario Squared Distance with Leave-One-Out (sqloo)
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a69ea086-f250-4866-94bb-414bd814c046 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment WPO: Enhancing RLHF with Weighted Preference Optimization
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6527333-4fa8-4f80-8b2e-aa2ad2bf0918 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c759edbc-5ff6-4263-80e9-adbb94926641 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Self- play fine-tuning converts weak language models to strong language models
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7551adda-81c8-4e25-a0e7-ed146cf186c2 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Evaluating Large Language Models Trained on Code
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d1c810-93ec-4e19-b669-115200e9cb7e · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58aef049-732a-4b19-bbe3-52df308e82d4 · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202baf4c-ee47-40ec-a812-86f42ce925fb · outbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Direct Language Model Alignment from Online AI Feedback
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af1d7231-a962-4122-a1fb-96b00ac11ab6 · inbound
The Differences Between Direct Alignment Algorithms are a Blur Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a01f68e3-aa5b-46ee-b15d-5caafd728814 · inbound
Multiplayer Nash Preference Optimization Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5af4bffb-29f1-4d3a-8af2-9b417ce00858 · inbound
NVIDIA Nemotron 3: Efficient and Open Intelligence Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.