Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2402.04792.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:27:11.442396Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:09:46.394396Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9fc03707-b6ce-471f-84eb-4d75de5d9fed · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Direct Language Model Alignment from Online AI Feedback
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 35a8d537-a80e-4db4-a8d6-ec6e39ef4f5e · inbound
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238780b0-f35b-4771-8497-9b696b970edf · inbound
Process Reward Models for LLM Agents: Practical Framework and Directions Direct Language Model Alignment from Online AI Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5724ed26-706e-4021-a240-b524f4676b17 · inbound
Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d71fb7-6b99-405a-9bc9-81f4537e1e12 · inbound
Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9cb827c2-9926-4f6e-8717-aabe5f2532a0 · inbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37779663-14c7-4844-8c03-42accb2ae545 · inbound
MPO: Multilingual Safety Alignment via Reward Gap Optimization Direct Language Model Alignment from Online AI Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec77987-1d39-4229-8a6e-f57c7ae1920e · inbound
Self-Training Large Language Models with Confident Reasoning Direct Language Model Alignment from Online AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · inbound
Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ae6733-b1b2-4dfb-88e2-ee7e74a226c5 · inbound
MOSLIM:Align with diverse preferences in prompts through reward classification Direct Language Model Alignment from Online AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1ca33b-face-4559-ab79-aec833b99512 · inbound
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Direct Language Model Alignment from Online AI Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f9ea8d-a73f-4b8a-a799-0a9968480331 · inbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Direct Language Model Alignment from Online AI Feedback
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258e424f-f7ff-40c9-a1f4-26cfa3be0f3d · inbound
Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d66ed61-f61e-490b-ae51-adb4cca2b999 · inbound
Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling Direct Language Model Alignment from Online AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e272212-0a67-482e-b5a1-9831005a09d1 · inbound
Customizing Speech Recognition Model with Large Language Model Feedback Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f5acc2-ba3b-43e3-946b-d31459bd35a3 · inbound
Bridging Offline and Online Reinforcement Learning for LLMs Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a140d1b5-77dc-4353-9028-07de6ffee18f · inbound
Data Diversification Methods In Alignment Enhance Math Performance In LLMs Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f54acb1-611e-49c7-ba95-7ebfe0a7f823 · inbound
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Direct Language Model Alignment from Online AI Feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d796d31d-9f3b-4109-8e94-f04c9c540e35 · inbound
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Direct Language Model Alignment from Online AI Feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e45fe2a-0649-4a2e-88a4-cf1875fbf5d2 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Direct Language Model Alignment from Online AI Feedback
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e3b838-9524-4038-8924-5fd067ad16f7 · inbound
Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7012f92b-6880-4eea-b3e9-6e20e0838a5f · inbound
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Direct Language Model Alignment from Online AI Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1431a99-d1bd-46e6-a6c6-9c58c7d34fcd · inbound
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation Direct Language Model Alignment from Online AI Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a917cb65-00e6-4af7-8631-da3807faf630 · inbound
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Direct Language Model Alignment from Online AI Feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50f7b995-cb23-4176-9e28-933b1a7e0e1e · inbound
Safety Alignment of LMs via Non-cooperative Games Direct Language Model Alignment from Online AI Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f530cf-96ab-414d-bebf-65dc1657cdfd · inbound
Ministral 3 Direct Language Model Alignment from Online AI Feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2215a8c4-a41f-4730-822e-34c9055982d1 · inbound
Noisy Pairwise-Comparison Random Search for Smooth Nonconvex Optimization Direct Language Model Alignment from Online AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c503c54-32c4-4024-a4e6-2084b118b98c · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33020fae-9d6d-41e6-aa54-059289939857 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c194993b-207c-4c12-8995-6db66a858ee6 · inbound
Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs Direct Language Model Alignment from Online AI Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07d33688-9cad-4a73-bed4-c96ca2a55bfb · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Direct Language Model Alignment from Online AI Feedback
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99e214aa-2fbe-46cd-a9d3-90077f9be17e · inbound
CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5777ec4c-c515-4c3b-8017-c2c5d5c9c181 · inbound
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models Direct Language Model Alignment from Online AI Feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40b92d08-c052-432c-b463-ba6f952ed9bb · inbound
Mind the Gap: Structure-Aware Consistency in Preference Learning Direct Language Model Alignment from Online AI Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6efc0a37-4cae-4f1e-8eb2-e890ecd24a08 · inbound
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Direct Language Model Alignment from Online AI Feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d527bc9d-9315-4726-ab55-5629f13f5ac3 · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a5303a5-7923-4d03-9a3c-e3cf9e802322 · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Direct Language Model Alignment from Online AI Feedback
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 485bf674-827a-4d27-9811-e11f5725fa08 · inbound
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Direct Language Model Alignment from Online AI Feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0cfec670-b681-4450-bb79-82a96b6e339a · inbound
dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models Direct Language Model Alignment from Online AI Feedback
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bcc9ea60-ea4a-4e1d-8449-e318756e9fb0 · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Direct Language Model Alignment from Online AI Feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dec601cd-24b2-4258-86b4-a06000049b14 · inbound
VSPO: Vector-Steered Policy Optimization for Behavioral Control Direct Language Model Alignment from Online AI Feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3a6c154-8285-4758-8d63-6191ecf29af2 · inbound
TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Direct Language Model Alignment from Online AI Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6bf4218-d171-4b0f-8898-8565ca3ab6ad · inbound
CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 303e8719-2e68-4ce7-a158-36d186402e70 · inbound
Trust Region On-Policy Distillation Direct Language Model Alignment from Online AI Feedback
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e34989b-6faf-4738-b3ee-e8305e6abe1c · inbound
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization Direct Language Model Alignment from Online AI Feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3fbc77f9-7207-494e-bd86-7239b59ed9bd · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Direct Language Model Alignment from Online AI Feedback
Reference 219
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5cb7cf2e-02e5-44ff-884a-07d8e2198004 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Direct Language Model Alignment from Online AI Feedback
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6fb746f-6fae-42ad-84f8-f9b65d2ac202 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Direct Language Model Alignment from Online AI Feedback
Reference 195
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd29c9fa-d3eb-4312-9572-46e39fe177cc · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Direct Language Model Alignment from Online AI Feedback
Reference 156
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459a5410-667f-4eaf-9155-22e366d49116 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Direct Language Model Alignment from Online AI Feedback
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd08babf-de01-4076-b7cf-6777e4df8941 · inbound
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding Direct Language Model Alignment from Online AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.