Pith. sign in

Paper Citation Record · LEDGER

Provably Robust DPO: Aligning Language Models with Noisy Feedback

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2403.00409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.00409 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:03:47.464904Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.406911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2a42b6b5-f3ad-4eb1-96b3-e897de916829 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.252104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:aefdb0cebd44e958117c44b00790e25307f9964946b832017cb254717668fbf5

Observation 0c4164c7-dd8d-4467-a398-af8bbeda1ebf · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.679899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:a36915fcb88e3c6eb918ee9bb5531a46cc9196b92c3f322b19c656749f64e569

Observation 6e9ef6e6-ad4d-4ddd-ac35-445bbf735d62 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.795434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.795434Z digest=sha256:03a1c855755da14032ca80517a3650cf23fb97c5d653517c302e965af68e808e

Observation 6f6217b9-2b21-4366-b350-b0d52c2f1c40 · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:08.819245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:08.819245Z digest=sha256:aa8a2f4a7b648ab0cae485418a0245477585774e0bc1fcc4220ee9d73948a9c1

Observation 55493028-8900-4233-ab07-3174d5ea4748 · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.320740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:fb1454a5e45d417be8624cf9d9d6a8c8fdf3ca2216627e74a654ae53f6aaf1bc

Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.673138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.673138Z digest=sha256:9b62489023e1c9480cb732f954f1d26b6aa9bf5ae5c67e1893974d31d27406aa

Observation 18383f47-ebb8-49be-82e0-db1deea0be98 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:53.423556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:53.423556Z digest=sha256:d983a63ac906aacab74d10871e130984cb1a7079a454bcf0c832d964ea0b6ad5

Observation e72950c1-3456-4bd4-8789-0c0fb597fb4a · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:31.844690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:31.844690Z digest=sha256:afaebbbebc4927996d95a1b645be238f4357d1c816fb7e010a6bc7b5eb46d557

Observation 879ddf9d-ee39-4e6d-bd23-4b8d4c93c977 · inbound

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates cites this paper.

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:01:34.180564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:58:26.626172Z digest=sha256:fa109b2913cc35ae9a7d428ec3179194d4cf5b7516590aac2a69e99ce7b9d5bf

Observation a0f62a07-09a5-43fe-8a7f-b05f5f17f77f · inbound

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization cites this paper.

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.570045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:27:22.570045Z digest=sha256:94421cac664f6e8f5ec6a0137b783caad6f3d194727b04036b6f6515eee33f40

Observation b1db7cef-8d31-405d-91c7-233f62db05e9 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.992814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:d951cb2d00dfb3919167cba45502fc93d59dda2af9a98fa11e0db88286d5e1a4

Observation 1f275270-dd03-4b92-9031-d2d2159fd45e · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.410770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T19:35:56.059362Z digest=sha256:9d8d99a45b5e72a451d4401b9413d1243233fc8eec6c9d5bad5f54b24b87e384

Observation 10a79da6-4279-4b68-9ab8-028dec443f34 · inbound

Multilingual Safety Alignment via Self-Distillation cites this paper.

Multilingual Safety Alignment via Self-Distillation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:41:00.754016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:08:32.264867Z digest=sha256:675554d43d8417484183dea07a1884a9fb9802da8f8f658a40a20d74e51fa471

Observation 59f00d15-ee8a-4cfe-8867-7658461bc1a0 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.273715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:c06834f16cfb390a589663a26c2bbb576e5886f81472f5a304a0277e11507a1c

Observation a29d9f3b-da26-431f-b3c0-00d7387e32e4 · inbound

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization cites this paper.

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:45:17.586731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T03:41:52.859647Z digest=sha256:1d3d9d78278b3c3636fa88a4211d549371096738f5023190b7300c3f11802e29

Observation 66a041cc-61ca-4ecb-90bd-fc72154bdac8 · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.254368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:adbb90719a40c9b0e218d6668a8019ce227eeee9fa44218ffe47a538b3ceaf7b

Observation 0d736ac7-41f0-4a83-9703-22e57c8659ea · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.408648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:53695620c780afa5f436b5a29f27c78a762bd00cfd0ed25d4595381cae65f00d

Observation 23266883-97f4-4d2f-8f72-348a70fd816d · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.410065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.410065Z digest=sha256:2673747e52e02981e2a1c52ae96c0703381a940c796ea0f519a0752e066ee8b9

Observation f903dc6c-5203-4d51-8355-c19cb17ee593 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:caa51a92f9e1e8a9b0f1ef543c48bdcb4175f36a828070915bf390f353650387

Observation 1b17ee32-4086-4ec1-9545-b8151cf549fd · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.937134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.937134Z digest=sha256:d16fd50a99bf098542bc29508ff6d40d5f84056caf24ba5aa25b1739ffa98d74

Observation bff2af43-1386-4c93-8032-bc0d540df63a · inbound

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation cites this paper.

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:03:47.464904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:03:47.464904Z digest=sha256:12416cc06d04eb5e442d3e9773ef15765943f40864db08dac7a811346f80507f