Pith. sign in

Paper Citation Record · LEDGER

Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.09279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09279 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:16.819602Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe69fe34-0db8-4220-8713-8b6916409610 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 182

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:30:02.969419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:67f6d2a004835a651bcd175cc83dc2f17e6442e2c227c2bae90047ef5e0c98ce

Observation 7e660493-93ab-42cd-a656-900de269cc29 · inbound

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective cites this paper.

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:16.819602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:16.819602Z digest=sha256:c1af1debcddd8e3b88d53f6c2d2105fc69378ee45102a05f490a52b566ac31fa

Observation ea815533-e427-4c7e-b099-20514c7aea49 · inbound

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs cites this paper.

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:29.551925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:08:29.551925Z digest=sha256:e409a58611162d31b60481c14a8d96624aa5fa3d33b75b66f4c6004db8631c7b

Observation dec2d96c-b540-40d5-b2fd-6ac2c13d463b · inbound

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment cites this paper.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.089108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.089108Z digest=sha256:886282e4d140cf3a931be50874a65c7bbb66ff4e4025eaacf7e7428c4ed10ae9

Observation 028f2908-3e8d-4bfa-a97f-5e506b9f1667 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.048128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.048128Z digest=sha256:1a2c53215f5eb54664c96b190871867b1b699e4181781b07b294951bb0621b30

Observation c0b7d5e2-33de-489d-9f63-86ba811a6a3c · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:22.272855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:22.272855Z digest=sha256:1299306517aa9a0b1b06b20cfe1da48c2552ebf3623ef7a67dbd91c76fe9fc3b

Observation af047bad-18ff-499a-a601-9678bb2bd54c · inbound

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality cites this paper.

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T15:38:44.895600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:38:44.895600Z digest=sha256:e4153423869fa15b49e3464bcd8eeeea081cc2c224b1d9e9939d66390f3e4cbf

Observation 3a882cb0-f77d-49a1-89a6-40eadf2bf368 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.172382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:4dcfc0797cf2824559043651661f80512a349c578d3ef8a357ee44c498df4f42

Observation bf82f9d8-9dbe-47c0-8da7-6fb0ce028d69 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.092443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:9cb4b12668cd6814679d98587d31ad3e5c71604c0c61ae520917c7d94418d0d9

Observation f4269c14-d1a7-41e4-ada1-f407f23400bf · inbound

Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion cites this paper.

Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:40.691816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:59:38.981808Z digest=sha256:c5134f51cb6d6e02f3b9f78a303435ac873cc0e6961aa0d7a39e79673f632f4c