Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2310.12036.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:59.145850Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation b045bfd9-05d4-4f25-8c19-72a780e85c46 · inbound
ORPO: Monolithic Preference Optimization without Reference Model A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 972de14a-91e6-49d4-94cb-6ccdaf5c574b · inbound
Process Reinforcement through Implicit Rewards A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation b3153567-2340-453f-ad62-cc1375bc3e29 · inbound
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a2c4d8-ace3-40d0-979d-4e89e722de36 · inbound
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f93d914f-0c4c-473a-82e8-2c013645e3ae · inbound
Thompson Sampling in Online RLHF with General Function Approximation A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9abefb6-8aaf-4792-96b9-8cd2dca89181 · inbound
Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a7ec8d-1f7a-4786-babf-f21180eb9c63 · inbound
Reinforce LLM Reasoning through Multi-Agent Reflection A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c59bc0-5896-4dd9-aef2-605091779e68 · inbound
DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1636fe-aa28-4d0a-a15d-841dcc051593 · inbound
CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1575096-4be4-46e7-b026-bd9c440d36d0 · inbound
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation fec1937c-8fd6-4e10-b526-f01cd2d18c30 · inbound
Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7e3696-975a-4dd5-9d5b-d36ed9ffd8d3 · inbound
Failure Modes of Maximum Entropy RLHF A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 4deb9cf5-bfb4-4a4e-a885-a13706c67041 · inbound
Adaptive Margin RLHF via Preference over Preferences A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d058d0b3-5f42-4060-9235-c76e24e0b6db · inbound
Safety Alignment of LMs via Non-cooperative Games A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03710aaf-d50c-4e56-b8a4-670bcc2e4707 · inbound
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation ae04bdcb-ef5f-4f03-9cd2-af5fee1a1754 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 2deb40d5-6e9a-486a-9ba4-13dc4827f5c1 · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 36ad412e-a310-41da-b084-7a7d3f8d7d7d · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 89c015fc-013f-4f68-85d5-8f127abb6416 · inbound
YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 4b7e7949-a768-47e2-85a4-642bbdd8a483 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 09a58498-7564-4110-a7c0-603eedba08a4 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation ce4112a6-da47-454b-a8cf-6dbba167bf8f · inbound
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 735f9d5d-4baf-4dcc-be85-c42485d78836 · inbound
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 5b41e8e6-8bf3-4891-842a-8fd6f1a8a774 · inbound
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation bb7add3f-651c-4f56-bc78-9cd92d032cec · inbound
Constitutional On-Policy Safe Distillation A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 75c04942-f7a4-49e3-81b1-6bda7cfbccb3 · inbound
FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 5b049f08-6e51-467c-9c92-342a0f7eb1a2 · inbound
Weight-Space Geometry of Offline Reasoning Training A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation a11da4f5-e823-4ed4-a425-56b5519e8320 · inbound
Towards Spec Learning: Inference-Time Alignment from Preference Pairs A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation 38364724-f099-488a-a77c-71f5162f8b76 · inbound
Towards Spec Learning: Inference-Time Alignment from Preference Pairs A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation a1437ddd-5690-4ed6-a8bc-fa94eab0e2ef · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T21:08:08.42545+00:00.
Observation b6fc0ef8-98a3-4acf-9eb1-8386c9e100f9 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e8bc7a-5eba-4e50-a3b7-20421d5bdd65 · inbound
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8fe755a-582d-4e88-b4ca-db7b229fe51b · inbound
Multi-Turn On-Policy Distillation with Prefix Replay A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 163
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b839b03e-a437-4a0d-a884-49efdc099734 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 164
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8899f132-d89f-4e05-bac8-94441a09aef5 · inbound
(Towards) Scalable Reliable Automated Evaluation with Large Language Models A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a51d4882-beae-4871-b73f-73d1a086b0f0 · inbound
Quo Vadis, World Modeling? A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.