Pith. sign in

Paper Citation Record · LEDGER

Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.16240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.16240 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:39:45.735113Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c765219-aed3-4e29-b9f8-1d133a61dde1 · inbound

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization cites this paper.

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:46.880609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T19:03:53.675107Z digest=sha256:b8222b19a8a592009281a09212772ca348179ce56e39b47a8b34a17d80d34194

Observation 4e4ee5fd-a9b9-4f08-95c6-cad6428fcaa1 · inbound

Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization cites this paper.

Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:39:45.735113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:39:45.735113Z digest=sha256:95fa924dc5013b9766f5eaad445b54df5b2a16cae1d4b285228a6a0da6be9164

Observation fe761636-83cd-4a26-a6ab-06a340bb0116 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:31.734703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:31.734703Z digest=sha256:10725a73926d4be2fdd357f56a74d350983325b75101f3b0e3751c19cad20eb9

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:73321c71214fd75693d45a2bbc78afbee3b5508e1e7dd8d38f4ba28b0126dfdb

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:b94088c0d47b02b9b067d18b88e2ab42cc7734a4b13b41f3c120d4a81e132418

Observation 16949879-736a-40a2-b1a9-74e0f343206d · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.039547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:87163f88c084792d358a1c1752684248a5f8430333180732481d3dbfdec00bfc

Observation 951663bd-9add-4f69-9910-0ce7a4bcc7e2 · inbound

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning cites this paper.

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.818786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:27:22.818786Z digest=sha256:14c317c8e2bbd09ee4650b4a8b27c707ce9307e775d73da674710bed72d898d9

Observation dfb966f1-9139-41de-8afa-bb2b116a9663 · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:36.093027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:36.093027Z digest=sha256:f87ef12baaf4b5f6f7661e7bcc6effc0632dde162883e103c8bd429e670e8a2e

Observation 1fddb961-0819-47c9-b429-8a4ed6fc1a58 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.126851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:5565ce1f061084e424a620cade367da48350fd614c309aa9b1ad353dfc9a7f8d

Observation 19973347-87d4-450a-83c0-7911744b8fe1 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.110302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:ac9165a8ed3dff2035fbbac3b99b2f2a68072bb570c5fb2988e4cb14e20ba5d4

Observation ada1bef8-3067-4a22-9776-478fcd096770 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:46:10.397358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:6734ce8260ed207a9e7bd76317833bba4e696d6d625d0b902db883264ca7426b

Observation 326e914d-0920-427f-9129-5c084eb5195f · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:04:04.467772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:3d609cce939ac82fcd41c9ea548dcb723bb0f4f348cbcb1d9e7f2bcfc241a944

Observation c695d87b-dea2-4534-a79a-71f609a6356b · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:58.803821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:87c3bb5cf5b1aff9775a8f324dca15620001b9652ef1939fe13ec7117c8aecc6

Observation 220a304e-7a5e-4c27-91f4-1e32b9584b29 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.019871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:6b0522580de028733e0c4b2373a85bfea004bb481a39fa1a05e71b9e4a01117a

Observation e0144991-53b8-4943-8b53-df773b4bd54a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.955425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:9255f85f9a2a71f4f0314c30f0d314fa2254fc7fb7f606b0c985442bc1e18308

Observation 0ba99aa1-1700-4fda-b7cc-82e8faffc46a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:e224638466f87141d506e97db264c3a66bab5caec901f978ca8b43b4d13907d5

Observation 243a7812-7707-4f01-a0f6-c10371f33d4f · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.245928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:705ef892a067aedf01eb09a6429642e5eb96516b2d9c605a34dbea53ff7944cf

Observation c9f3e4ec-e3a0-4560-9122-0cf4de55d24c · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.415473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:34eb8ff4cceb08237a922e539160b828a2a9770d0d8f9a2e7b807ad07d5d62d2

Observation 6c795686-ad62-4af4-85df-7b8a941c1ae7 · inbound

Rethinking the Role of Temperature in Large Language Model Distillation cites this paper.

Rethinking the Role of Temperature in Large Language Model Distillation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:50.064284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T23:25:55.849389Z digest=sha256:91a7eeaf32a89b42f1f9143f5cd8f0f8533a616bb5e0cd471139c9f8f08df9a7

Observation 3c744dee-790b-4f01-b78a-8ec48199d654 · inbound

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter cites this paper.

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:56.186095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:53:22.159132Z digest=sha256:7c2ce5fb435a282b83259f125aca9af996958b4762c24beb9ed65f55cc17eb3c

Observation d976553e-09a8-4127-a434-1fbf0a4f9173 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.380972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:555d086b72f1bde2bad594fa75fe983de72b426316e721b24a9e03d424a9ba66

Observation c6053d29-2b05-4f20-9194-e5acdafbac21 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.407194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.407194Z digest=sha256:a7939b07ad6f4c86dc1d61bc9401ccf1f91fbe89c90026e17c964d927c5dbe9c

Observation ff8e3dd2-af1b-4ce2-b914-fb03b68635aa · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:de1e6803182a451e00ee463646e0500e79dadbebca3faaf90710aac221f4dd1a

Observation 1d1df6fb-b34e-4886-bc4b-e260e9ca3516 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:50.233277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:50.233277Z digest=sha256:8ae3db4647c78e412f54a2fda9b3218c9505d16c12b607bb1ffa990909c69c62

Observation c8fc713c-d568-4789-a4ca-d889d967af12 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.977272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.977272Z digest=sha256:a66d76e0776618aaede7dfcc8180b06928bb18036ded16e4164aaab164c2faa0