Pith. sign in

Paper Citation Record · LEDGER

AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.10148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10148 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:32.415979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:57:48.045665Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e2f53a85-0537-4637-9bee-cbda48e618e1 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:32.415979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:32.415979Z digest=sha256:f41861a7b321b4b74e790789f22153cb5c64c325c9b5a2692851a32f2aaaf602

Observation 9402494a-17ff-40a8-a70f-fee6e656ded9 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.506011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.506011Z digest=sha256:e9699a46472d8cd1e5a4dcea4ba64a995cf537a25733b66171292ef974cfab08

Observation 5a83de81-fdd1-495f-9879-139ee7e7b6f3 · inbound

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization cites this paper.

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.596891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:27:22.596891Z digest=sha256:8b32a44bc81343c4238061cbf98b8f6dff214dd682c5e7be62728aefbb213437

Observation 9d29c634-d292-4171-8398-95cff374d276 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.538427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:7519a07ce9825b2e63df823e8a90c6ae7de005d735f23ea77d2a91d429375142

Observation 3e9f1084-dab5-4e3f-8144-652fbcb8251b · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.403397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:e98c80aad2817b8d22f2340ff3490f62811c8b9a2299fd45e00f828e64ea7578

Observation e864625f-9dcb-4216-8adb-e219b9fce24f · inbound

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training cites this paper.

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:08.749089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:17:57.772960Z digest=sha256:2bb4b0a430cbff3c91868ef5f7f5b65c48c2ef53c88c72d96aec08888b2a2e07

Observation 82765739-c562-49e5-a9ec-8391891e004c · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.570076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:c2b0d11196aa7fa91da060d5f3e833b4f7620785a8cc57c10f4e77e312abffc2

Observation 811948aa-c2cf-4208-b272-d08ca682b388 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.835198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:14b1f1fac7b2db39010b3ec8a3334d78a953ab89d49068fcdb224ce6d4e1a1c6

Observation 8015c34e-4fe0-4272-ac4e-ea1f9bb5920f · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.722496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:e5b3648d6f1a77dd22c5cd40167d77b107127b0000b5af8058654b07d38f5a66

Observation 2a964291-e06a-4f14-83ae-f4d9cc84d058 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.912232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:fe4c95f7cbdcf4375296b7edd7679839ddd2f5572cc901242b9be26c5441c2bc

Observation 9855c0e9-c1f1-4831-9e5c-8ef75d1794a7 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.750597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:0ab9e151a424a4c0c210467b069236243c29b73aa02ab4b6a2925dbbcaff16bd

Observation e45fb5f7-635f-49fb-b37f-e5cfc184354a · inbound

Boosting Direct Preference Optimization with Penalization cites this paper.

Boosting Direct Preference Optimization with Penalization AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.047259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:36:26.784333Z digest=sha256:07de988777194aa211b410474ae016fa7f9f892abdd979ddd88cb2d10740979f