Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2409.03650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:56:28.735011Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T13:49:14.274811Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3e3f52cc-106b-49cc-a082-07622696c53c · inbound
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6ea5a6-77d4-4dbc-9a9f-579e7aa307f3 · inbound
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f0ace7-0ead-474d-aaaf-3a7fb6b1e2fe · inbound
Explicit Preference Optimization: No Need for an Implicit Reward Model On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90df85d0-0cb4-48ec-a842-12ca096028a7 · inbound
A Survey on Large Language Models for Mathematical Reasoning On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9e3cb7-3729-47c1-9aab-1d4e935edfe0 · inbound
SGPO: Self-Generated Preference Optimization based on Self-Improver On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.