Pith. sign in

Paper Citation Record · LEDGER

Provably Robust DPO: Aligning Language Models with Noisy Feedback

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2403.00409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.00409 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:03:47.464904Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.406911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2a42b6b5-f3ad-4eb1-96b3-e897de916829 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.252104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:c835e79fffde5022912b518dc957e7b594ba4f8ee99d6f229ae8a5bedc2ef998

Observation 0c4164c7-dd8d-4467-a398-af8bbeda1ebf · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.679899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:21a24d629af60baccbd875d3af663e62b6d62da1b530e682ca23daea4e9a2e90

Observation 6e9ef6e6-ad4d-4ddd-ac35-445bbf735d62 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.795434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.795434Z digest=sha256:c3631ce06d8d39c0459c9b642d419716e75e9c912d8417d08e9676af06a4dfdd

Observation 6f6217b9-2b21-4366-b350-b0d52c2f1c40 · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:08.819245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:08.819245Z digest=sha256:66bf57596ba811516d2a77cdba268b46c3304257311405a50a42f0e663f90dad

Observation 55493028-8900-4233-ab07-3174d5ea4748 · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.320740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:bb3fc448d5af8edc9d1be4b8e03329ed4586841d378a360db460b465aed9956e

Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.673138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.673138Z digest=sha256:11eccaffbfc2d30ae21c45824717041f3977bbde81990fd7d5d01a2fcb04a1a1

Observation 18383f47-ebb8-49be-82e0-db1deea0be98 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:53.423556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:53.423556Z digest=sha256:775e4465a0a1659e097e8316ae8e4c16a28ee417fdf9173a2cbe2c82b6b939b5

Observation e72950c1-3456-4bd4-8789-0c0fb597fb4a · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:31.844690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:31.844690Z digest=sha256:19afcdca73ed6fc39e21d3d47ff0e0629789201f6ca543eb8794a0662f8d650d

Observation 879ddf9d-ee39-4e6d-bd23-4b8d4c93c977 · inbound

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates cites this paper.

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:01:34.180564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:58:26.626172Z digest=sha256:335e51e46e7c169c4f5b8bbd3d30addf4fd8b27715a67b7f8e088db80b94602b

Observation a0f62a07-09a5-43fe-8a7f-b05f5f17f77f · inbound

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization cites this paper.

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.570045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:27:22.570045Z digest=sha256:8bda5422faa7e1903018a49f4523a76582841b4a2130b0a880ac1ea965998060

Observation b1db7cef-8d31-405d-91c7-233f62db05e9 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.992814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:8398009ae993437d05f621c3602aa151283ba79bd04e5ebd49dca67a73ad81c8

Observation 1f275270-dd03-4b92-9031-d2d2159fd45e · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.410770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T19:35:56.059362Z digest=sha256:3900b3879456ed9b65bdd5c76ab6f85936d25526689d3bbee327dfc23020aa62

Observation 10a79da6-4279-4b68-9ab8-028dec443f34 · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:41:00.754016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:08:32.264867Z digest=sha256:1af1602409eba9ca2f11fbac30009d88462c315b44d4e34f1961b74d3d6e9200

Observation 59f00d15-ee8a-4cfe-8867-7658461bc1a0 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.273715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:abd56dcae879d808d7f941b3c3f2917de45a3f08ab5ab9cf498651b3d180f77f

Observation a29d9f3b-da26-431f-b3c0-00d7387e32e4 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.586731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:ffb08aea07e78d75c8cf5107d38853b09afd4c2a3f1970f093c1237f236dcc5a

Observation 66a041cc-61ca-4ecb-90bd-fc72154bdac8 · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.254368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:0b69f1b954c358804fbe6b1e6c2d146f3b8a1dab8d9ff3c7e5168723e8a27bd8

Observation 0d736ac7-41f0-4a83-9703-22e57c8659ea · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.408648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:2009894a6fb7bb55922fc2d739254409e71686a289935457a2e55fe1471cb549

Observation 23266883-97f4-4d2f-8f72-348a70fd816d · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.410065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.410065Z digest=sha256:e84d17323f19c2795ea4a50043e814d822ee45ea37fc5c4b38f9e6414208985c

Observation f903dc6c-5203-4d51-8355-c19cb17ee593 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:8c7e980948ffef20fbdb82cfc30c025dffe54b01271ffb43312d7d1cd7b6b0d4

Observation 1b17ee32-4086-4ec1-9545-b8151cf549fd · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.937134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.937134Z digest=sha256:de1cc36872e36fe9c2bfe6f8c727ee596a5a2f3f5597608a37f4273b42be7579

Observation bff2af43-1386-4c93-8032-bc0d540df63a · inbound

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation cites this paper.

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:03:47.464904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:03:47.464904Z digest=sha256:9cf150832b50ee518e47e700f59f51d72d9ea658edef3686e191d76831e9a3c3