Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2410.15595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.15595 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:41:20.729703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 519dfa99-5557-425f-8788-51f51b6b75ce · inbound

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts cites this paper.

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-23T08:00:12.781392Z digest=sha256:8b67dc793d3b29abe1dcdd1c1500a39a682a06909fdabfb5c7905e855b188dbb

Observation f028a299-d28d-459f-84e3-0fa6a25f4242 · inbound

A Systematic Examination of Preference Learning through the Lens of Instruction-Following cites this paper.

A Systematic Examination of Preference Learning through the Lens of Instruction-Following A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:41:20.729703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:41:20.729703Z digest=sha256:a7fc4932d0b520dbe46ed508beea4e72ee141a5ef08bdb8751a59716fc83e2f1

Observation 0d3620bb-1f7d-453e-98d8-255766d04571 · inbound

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems cites this paper.

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-10T22:51:52.728580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:51:52.728580Z digest=sha256:0179782db65ed31271c02264b77b7dde8bde051b296850fec7039c0780d0411b

Observation 5245b116-1829-4987-b49d-f4d9e617ab9a · inbound

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information cites this paper.

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:56.533627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:56.533627Z digest=sha256:4e13d689f9f58ade128d96129855af6f4bcd8dd974172afc0508a7a5b69b9b83

Observation 34e2bed6-a7a6-4cbe-bce3-07af2c1df1bc · inbound

Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization cites this paper.

Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:22.245835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:22.245835Z digest=sha256:94985a1acb69f7bb9f0750d965165a2bf146c2b5053f0e06523224c7abec9104

Observation 771e57d2-c34f-4158-9bea-d118b007f221 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:23.510307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:23.510307Z digest=sha256:bed50b2774a97d2b2690986d1c93432ed87eebddc8b5da7c38165221649819b0

Observation 9e4f5b25-ea62-4f09-bbf0-0c4acd5eda14 · inbound

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations cites this paper.

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-19T06:09:26.269452Z digest=sha256:83eb0497cbde0f47217e1187b432539a7ad1821833e0a66a145c28748f7204a6

Observation 53a7ab6b-16a0-4246-bd5b-83e320dda34c · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:705ebedfa5b5625f5cbb01964963d2ec0ab1a34137ec083853e5df7945dba16d

Observation 98c858db-038a-4c6e-9a94-6d9fc951bdf2 · inbound

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization cites this paper.

Active Causal Experimentalist (ACE): Learning Intervention Strategies via Direct Preference Optimization A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T05:26:01.675314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:26:01.675314Z digest=sha256:c50d419188c058cf79ceffb14aa02c331d5ab291cf7d3c90d5176496d4f7d6b7

Observation 3013383f-6f4a-4153-95b5-e41680157898 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:738486004f064d001adbc01c592c89667278fb4ef65924245aafc050f4c300c6

Observation 72b588ce-8bf1-44d0-8e97-596c09337a06 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:ca8cf3ea2b668b356bfc011b2057d083cbfef1c1fa882727e5edeaf1ad07c48c

Observation 942de73f-1e2a-400a-883e-7b98baa62f04 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:c54312cac1a22b8fc53d5ce12e9e885abe0a0efab94b882f67951272ec1919e6

Observation 02f081b9-9386-4801-9343-faad5808d4b8 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:a5fe8bcc2d4e1c9c595b118edbe9c3b9d0d326e88d7f36bd2c03efa0889a2d7c

Observation 5d75473b-9ac6-4a3d-b4b1-9e20570849a5 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:b3e1de068256c4d88bdc2b1f7c6cbcdd113b01828397c789850140317508ab24

Observation bbb00dc8-427e-4690-98d6-7d2477d6ca53 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:23e322283c671d01e83968f5b5f44f83742d349f2af566a4b2faabb8293d2e55

Observation 88539dbc-8626-4e97-ba06-9cded6ea9dac · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:10.714292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:668a6fd93e3aa9e0c54e4f908c3e3a620cb5422577e365e198aaf3af5aaf1815

Observation d2a656bc-9798-4349-a83e-5b4be92d3abf · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T22:25:06.439296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:9ddc40d2647673f7c069b4f7aca8122413d450875d02839dae7f77fce3634d4a

Observation ab5e5094-07ec-4c79-9063-00f45ac2f4b2 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T22:25:06.431552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T22:15:53.639048Z digest=sha256:860aecdcad8078d85a2442fa31efa6ead240f5523293fb5345bb158e0ce7327d

Observation 97fe8aa6-09c7-463d-80d0-c2b16d2dc714 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:19:02.875664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:d23926c909b86d8e2a8d7d076c6cec4c44a5ecb02b8ac2704e46831f45a26e8f

Observation e86b1d7e-f9db-4e1d-b288-cefa5e6e7f8a · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 233

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.699362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:9589d76b88cd00d8d0e80634de1c2223ccc4c9b38b017b7f96dbfb55eaa46db1

Observation 65e0b408-5a26-4e3f-ae87-b9879ef5963c · inbound

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design cites this paper.

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T10:59:18.707247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:59:18.707247Z digest=sha256:ff278d559e124edcf2ce746d00828c91ebfe9ea1ab0bf90cbbae2054ab5eac21