Pith. sign in

Paper Citation Record · LEDGER

Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2404.04626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.04626 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:50.883420Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:34:02.199643Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 38d3c96e-9742-47fd-a1c4-3bf2dea718c6 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:04:44.473510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:f49b7d573eef7ce77767d38825447909bce03901840e8a0fb88ed0f7cc9dcce8

Observation 6fb41568-fb5b-43c1-93ad-7c7ab2d28a3e · inbound

Continual SFT Matches Multimodal RLHF with Negative Supervision cites this paper.

Continual SFT Matches Multimodal RLHF with Negative Supervision Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.740552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.740552Z digest=sha256:825c307bfb534e4b79129652febafe3a7733039349452797030d50e56290642b

Observation 2e3a336b-d35e-412a-a212-196ae26a005e · inbound

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability cites this paper.

Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:43:31.415705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:43:31.415705Z digest=sha256:e7b35ed30a3396d25a03b07a683b9e72b390fa2e9348e42814a0cb8860bb6a4d

Observation 2040eab9-30f9-42ad-9471-7d63ba961cee · inbound

SPRec: Self-Play to Debias LLM-based Recommendation cites this paper.

SPRec: Self-Play to Debias LLM-based Recommendation Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:41.218272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:41.218272Z digest=sha256:9d02a57e2d4fab6735785f653d31f188a569c682d22283dd315b6dd63a4a19e2

Observation 0bf201f8-6681-4634-b04a-db8e7bd31ba7 · inbound

Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator cites this paper.

Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:11:09.137472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:11:09.137472Z digest=sha256:60f4c424845714aac712f6df6f2cedf4e4c267128e56d02a24e7f5d9abf815eb

Observation ea2b61fe-61ac-47a8-9026-a9ad5900318c · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.888937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.888937Z digest=sha256:ce8a18d2eab64bba2e95049714104fb70bb5e89a6025ad0d43f5657b91a537d0

Observation 1f7f027b-a5ae-4010-8ca5-2a6cf7e8ea4b · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:32.212732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:32.212732Z digest=sha256:589935afa2719d9e886972735da51efec011ae889845b4e11b0e4c90593ed1e4

Observation 7a277e6e-d56e-45de-8967-eb401de4836a · inbound

SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment cites this paper.

SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:50.883420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:40:50.883420Z digest=sha256:e97a60483aeb71482d4c8083a7fd54fb80f98b8d193777f0238cf4f6930f1d9c

Observation 6a490680-d938-4d17-8bea-71631ddfb1bd · inbound

Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy cites this paper.

Empowering LLMs in Task-Oriented Dialogues: A Domain-Independent Multi-Agent Framework and Fine-Tuning Strategy Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.943627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:08.943627Z digest=sha256:cf147486e9efae12a5280e4a889bb9f8bfed8de7f6afa843c5556cbaf6b04f3a

Observation cd7ded6e-a45e-4638-8e46-8fcd2e28457f · inbound

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists cites this paper.

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.669169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:41.669169Z digest=sha256:e7c0cbc26e935f950421e28d5ca7037d7a3f2f6a57bb46e56e1fdc723100cde4

Observation 29c4a37d-1f34-468f-90e3-68cb338f922a · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.882471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.882471Z digest=sha256:b8bff7ffd41f2ebe4930c8339b2a26826caa50d7562c457429a350ec173c18f5

Observation a3bbaa8d-cb41-42d7-96f4-f338b417b477 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.609799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:45.609799Z digest=sha256:8295b715b91cd500fc3f04f6671d24ac085cfadc874c3b99d07fc4cb096c73a0

Observation 11213979-cfc9-444d-a231-6100890f1019 · inbound

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents cites this paper.

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:11:18.299379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:11:18.299379Z digest=sha256:5b5eb496fff9d27510c4d5bbcf7f537ddb1c79ed02ea51b04774cebd748eed3a

Observation ff4fe65e-d1f5-4eb6-96dc-0c36871e3cd9 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.898254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.898254Z digest=sha256:b645fdec155f5cd88a59fed9ed7fa0241ac657de80555152b4dbe78919e318c3

Observation 75974a4b-ac77-4009-a2b2-a3c5a7634cfb · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:04.905986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:741c32310d6032af40b80acd523e4e8d75603b63b21b4f1a017733521c45074c

Observation 1dd85c43-ba10-43c4-871f-4441a212aeaf · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.062549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:556e38f1ad8e72be9b6b6f583b5d207a6396692254e35935b69d7b973a29ff6d

Observation e8247885-8341-46c1-b3e7-f8dca5f1ce15 · inbound

Curriculum Learning for Safety Alignment cites this paper.

Curriculum Learning for Safety Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.202851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:25:31.739331Z digest=sha256:5751dd4615c282047ea783b8aa6eed8e6a6a3fcf11ed5f2571bec15c93d9f604

Observation 6dec9c58-feae-4e56-91fb-9c479763bc1b · inbound

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates cites this paper.

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:33:24.405470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T12:29:55.729913Z digest=sha256:9de88c3766c1e2015c4032f118af81a5b6ff7ec58a77cb7c79c4815cea2d01a3