Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2309.16240.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:31.734703Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 0c765219-aed3-4e29-b9f8-1d133a61dde1 · inbound
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe761636-83cd-4a26-a6ab-06a340bb0116 · inbound
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · inbound
Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16949879-736a-40a2-b1a9-74e0f343206d · inbound
Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 951663bd-9add-4f69-9910-0ce7a4bcc7e2 · inbound
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb966f1-9139-41de-8afa-bb2b116a9663 · inbound
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fddb961-0819-47c9-b429-8a4ed6fc1a58 · inbound
Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19973347-87d4-450a-83c0-7911744b8fe1 · inbound
Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ada1bef8-3067-4a22-9776-478fcd096770 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 326e914d-0920-427f-9129-5c084eb5195f · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c695d87b-dea2-4534-a79a-71f609a6356b · inbound
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 220a304e-7a5e-4c27-91f4-1e32b9584b29 · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0144991-53b8-4943-8b53-df773b4bd54a · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ba99aa1-1700-4fda-b7cc-82e8faffc46a · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 243a7812-7707-4f01-a0f6-c10371f33d4f · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c9f3e4ec-e3a0-4560-9122-0cf4de55d24c · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c795686-ad62-4af4-85df-7b8a941c1ae7 · inbound
Rethinking the Role of Temperature in Large Language Model Distillation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c744dee-790b-4f01-b78a-8ec48199d654 · inbound
SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d976553e-09a8-4127-a434-1fbf0a4f9173 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6053d29-2b05-4f20-9194-e5acdafbac21 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 177
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8e3dd2-af1b-4ce2-b914-fb03b68635aa · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 158
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d1df6fb-b34e-4886-bc4b-e260e9ca3516 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 159
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8fc713c-d568-4789-a4ca-d889d967af12 · inbound
Normalized Rewards for Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.