Pith. sign in

Paper Citation Record · LEDGER

AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.10148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10148 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:39:22.659889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:57:48.045665Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a60a2d6f-201c-476f-8f32-95acc3e726bd · inbound

Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning cites this paper.

Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:39:22.659889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:39:22.659889Z digest=sha256:14d20b25b586272b7af33ba51e29935f6c06020d2baca01d6c247d0e34845155

Observation e2f53a85-0537-4637-9bee-cbda48e618e1 · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:32.415979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:32.415979Z digest=sha256:f41861a7b321b4b74e790789f22153cb5c64c325c9b5a2692851a32f2aaaf602

Observation 9402494a-17ff-40a8-a70f-fee6e656ded9 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.506011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.506011Z digest=sha256:ed576a06bf4e31ac8f17b9873ab659ce71aad45099131368980b394389101b1c

Observation 5a83de81-fdd1-495f-9879-139ee7e7b6f3 · inbound

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization cites this paper.

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:27:22.596891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:27:22.596891Z digest=sha256:8b32a44bc81343c4238061cbf98b8f6dff214dd682c5e7be62728aefbb213437

Observation 9d29c634-d292-4171-8398-95cff374d276 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.538427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:a4c00db4f3b0f437fac7959c51c3c9e6d9af0c5f79e57d9cec06644435fdb732

Observation 3e9f1084-dab5-4e3f-8144-652fbcb8251b · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.403397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:2acef54bed86e63c19ba0a8515bea30cf7a645533115306af01d0e020f7601b5

Observation e864625f-9dcb-4216-8adb-e219b9fce24f · inbound

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training cites this paper.

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:08.749089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T14:17:57.772960Z digest=sha256:d8f3c9e96732391671f79ca9ef6cf4ff87587cc0ca4bbd35073abbe30dd73110

Observation 82765739-c562-49e5-a9ec-8391891e004c · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.570076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:debfc8d5fe522614d7961ff958c9f8fdd4ac04b069d423680dc91219a349ffa1

Observation 811948aa-c2cf-4208-b272-d08ca682b388 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.835198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:33b90e7231356a79b189907662d832c62c19cb2df4dd47056f70fe2cbf66dc3c

Observation 8015c34e-4fe0-4272-ac4e-ea1f9bb5920f · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.722496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:b87eeed9ff5a92c4db4af6aa2941025f8c8c7a38a15fa0ffe1be035981ed8674

Observation 2a964291-e06a-4f14-83ae-f4d9cc84d058 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.912232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:a74a810a99e03faed41aba683d3e29c15cc058e56d1208a71645c140a01d4e88

Observation 9855c0e9-c1f1-4831-9e5c-8ef75d1794a7 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.750597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:e5392b7d494c869a51b0d532c5056d45750763efbe57043efde3bd343f0fed6c

Observation e45fb5f7-635f-49fb-b37f-e5cfc184354a · inbound

Boosting Direct Preference Optimization with Penalization cites this paper.

Boosting Direct Preference Optimization with Penalization AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.047259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:36:26.784333Z digest=sha256:bcdf24da2610aa970d68372d62818a56bde5e19559f9dda2590c3c98699d7e05