Pith. sign in

Paper Citation Record · LEDGER

Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2309.16240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.16240 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:31.734703Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c765219-aed3-4e29-b9f8-1d133a61dde1 · inbound

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization cites this paper.

Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:46.880609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T19:03:53.675107Z digest=sha256:f8dff332b74eb4027127ab6f63e1b9efe82e334849325dfad00e92ee728cdb83

Observation fe761636-83cd-4a26-a6ab-06a340bb0116 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:31.734703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:31.734703Z digest=sha256:ef5b856096c6e6155ad627b66f99aec49f8ebf8997d97d1c4187a6a42d5d932e

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:73321c71214fd75693d45a2bbc78afbee3b5508e1e7dd8d38f4ba28b0126dfdb

Observation 76980b50-eb90-47f5-9a42-dc7bc440db34 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:49.083529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:49.083529Z digest=sha256:fa0c15ea9b17a034ecbcbd85bb87353bc34b52a25f602462e46a8f448cbcf8e2

Observation 16949879-736a-40a2-b1a9-74e0f343206d · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:24.039547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:2777fcd642a86ce340d31a6d57771858d6888c137e95529c7884e38eda53be6a

Observation 951663bd-9add-4f69-9910-0ce7a4bcc7e2 · inbound

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning cites this paper.

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.818786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:27:22.818786Z digest=sha256:14c317c8e2bbd09ee4650b4a8b27c707ce9307e775d73da674710bed72d898d9

Observation dfb966f1-9139-41de-8afa-bb2b116a9663 · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:36.093027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:36.093027Z digest=sha256:f87ef12baaf4b5f6f7661e7bcc6effc0632dde162883e103c8bd429e670e8a2e

Observation 1fddb961-0819-47c9-b429-8a4ed6fc1a58 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.126851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:25dbc42093f776a236150b5ecc56d8bb195cd5275e20cb88fb741381179c1fe4

Observation 19973347-87d4-450a-83c0-7911744b8fe1 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.110302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:7a7dbc208b023091cc955e43374e7398c8c18540e2c74cc1cfd1d934132c9d8a

Observation ada1bef8-3067-4a22-9776-478fcd096770 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:46:10.397358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:2595ceceee693996104af47b1a315f03c8a2e6d5f97784debef30f2dc14642f1

Observation 326e914d-0920-427f-9129-5c084eb5195f · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T09:04:04.467772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:2111493a48d0da2251d851242e282f35d519ee3aca8330ff82fe7f176641f9dd

Observation c695d87b-dea2-4534-a79a-71f609a6356b · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:58.803821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:a57f238992f2306e2c2f0182f848a2b185dfff18298ccf8ad31b671cf0e31456

Observation 220a304e-7a5e-4c27-91f4-1e32b9584b29 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.019871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:08fa934a3da03e137f75ccacdcb4ff24e8863204a845497e299df640284df1f7

Observation e0144991-53b8-4943-8b53-df773b4bd54a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:27.955425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:6970b61417ea9da302b5427db435c093b2475d84d146a8d35d09e3fb43802381

Observation 0ba99aa1-1700-4fda-b7cc-82e8faffc46a · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:c61c9c444a4b529d10abe67cabbfb3e635764ad0568922a74bc977724b6cc0ac

Observation 243a7812-7707-4f01-a0f6-c10371f33d4f · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.245928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:63b202476b3022648012a89c4b9578c75e11733466e987447cfcda6a8067d141

Observation c9f3e4ec-e3a0-4560-9122-0cf4de55d24c · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.415473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:4e742c2c30f07acf7565dc1b800ba546023bcc77da9069afddb704f9d2f24e51

Observation 6c795686-ad62-4af4-85df-7b8a941c1ae7 · inbound

Rethinking the Role of Temperature in Large Language Model Distillation cites this paper.

Rethinking the Role of Temperature in Large Language Model Distillation Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:50.064284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:25:55.849389Z digest=sha256:dd69df38037017a63ba71b81aeb2e3c8f06dd92ec27744c4936f00b4f42a5c9c

Observation 3c744dee-790b-4f01-b78a-8ec48199d654 · inbound

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter cites this paper.

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:56.186095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:53:22.159132Z digest=sha256:69963cb5a2be02de55bd93c224a150be076af99e745dd15676fb80a6215bc278

Observation d976553e-09a8-4127-a434-1fbf0a4f9173 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.380972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:cb3d181a4b3ce96170a2e4f193a60033bdbbde1907a4a61b88c41301ef54c249

Observation c6053d29-2b05-4f20-9194-e5acdafbac21 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.407194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.407194Z digest=sha256:a7939b07ad6f4c86dc1d61bc9401ccf1f91fbe89c90026e17c964d927c5dbe9c

Observation ff8e3dd2-af1b-4ce2-b914-fb03b68635aa · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:de1e6803182a451e00ee463646e0500e79dadbebca3faaf90710aac221f4dd1a

Observation 1d1df6fb-b34e-4886-bc4b-e260e9ca3516 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:50.233277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:50.233277Z digest=sha256:8ae3db4647c78e412f54a2fda9b3218c9505d16c12b607bb1ffa990909c69c62

Observation c8fc713c-d568-4789-a4ca-d889d967af12 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.977272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.977272Z digest=sha256:a66d76e0776618aaede7dfcc8180b06928bb18036ded16e4164aaab164c2faa0