Pith. sign in

Paper Citation Record · LEDGER

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment

As of 21 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 2 inbound Pith citation observations for arXiv:2506.12725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12725 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:52:40.142142Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T10:33:54.851493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.077289Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eea05dbb-9d5e-4c5d-96c5-eeda1ea10ebb · outbound

This paper cites Nemotron-4 340B Technical Report.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Nemotron-4 340B Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:36.504800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:36.504800Z digest=sha256:342018fec1fa500b215211075ace98395c862c87008abd00ff7c204b989d96b6

Observation be3c1356-7b0c-41fe-b551-d2a62c7d1b27 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:36.583198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:36.583198Z digest=sha256:1ed1b1af1441d175d45f70324eabe9caeea8fe2982d5db4a708a1a4ac509391f

Observation e8416cc5-e960-429e-a412-41d57b747108 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:36.684782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:36.684782Z digest=sha256:800883c1bffe0fe631b758aeb8e6701504159ae717eb89daa7f1792bf8d9fc66

Observation fecc0cc1-0998-40c1-9b0a-4ca3c7c5e8cf · outbound

This paper cites Preference Learning Algorithms Do Not Learn Preference Rankings.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Preference Learning Algorithms Do Not Learn Preference Rankings

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:36.777313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:36.777313Z digest=sha256:6453df47148ebb5ff49b99afa54f466245eab60fda6406c5ad9a0a938e43469b

Observation 584d3c2f-5e80-4a77-a6f5-9f8fc7961a0c · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:36.897991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:36.897991Z digest=sha256:9d85aa1b661a99a7ea3b6e41e78d108650d7107ea593800e0beae3c2da22e9af

Observation 7e280702-b685-4bb0-9a7d-b0c727e6140d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:37.016549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:37.016549Z digest=sha256:aa1790028324a37382cd198a34c81fd3102901e2a6f00c4948eefd9b937a74aa

Observation e321a7a6-f36a-4d5d-b7c3-319fd6a1159b · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:41.861680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:37.102719Z digest=sha256:85d732745026ab2ad9ceb9be83680d42d8852aa2b3a34ecc2ab20f56aaf5fdc5

Observation 2626964b-c71b-4c81-8e7b-ebe65b5487c9 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:37.207357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:37.207357Z digest=sha256:2948e2a46be717898726fd55c0747c2d7f544ba547e578a4888a6aad44d24ddb

Observation 6758a168-2c5d-403b-b989-2374ebf5ca1c · outbound

This paper cites The Llama 3 Herd of Models.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:37.397258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:37.397258Z digest=sha256:d7e60ccb1d27bce742c5ba3d9f019303201db971af1773c3343495a8632b8868

Observation 69614674-8c3e-4a0c-b2a7-2212109711c7 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:41.832367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:37.528114Z digest=sha256:33ef90f07e5995a34ad2e0438a001e9bac9987fea9daf3d60d5f3fe3b2ed9eda

Observation c7658ce2-70e4-45a3-9b4b-da6d6aeb9b54 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:37.786943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:37.786943Z digest=sha256:61329b7970ed5e4bf171f71aff35b086593d22a481e1b2ed587d94befb20e41e

Observation b6e5fefe-5cb4-4e2a-927e-7b4fce63c90e · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:37.953089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:37.953089Z digest=sha256:7596eb23eea8dae7cc35ea2c88a6faa262d5034ef698a687373c1f152074eb07

Observation dec2d96c-b540-40d5-b2fd-6ac2c13d463b · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.089108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.089108Z digest=sha256:1fd5661cf1b6bd1e4b9e68965c33d6bcf201c5c05ca0c56e3602d7458d64d71c

Observation ae148c67-3aee-4709-bed9-cfe1c368eacc · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.224309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.224309Z digest=sha256:0b36305c27c52820bc9d3696ba12ae7945ad30d7910c854d3253b302209ea4bd

Observation 636ebf8b-c0cc-465f-8c3f-320acc1071a4 · outbound

This paper cites Decoupled Weight Decay Regularization.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Decoupled Weight Decay Regularization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.366472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.366472Z digest=sha256:b66d28ce28667cdf57901efad3033046b2fe36d21f5a6b0ff4f5c2a48ce7865e

Observation 52dc2c5a-4574-4702-97f6-6ae8a112963a · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.506533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.506533Z digest=sha256:26360c8effe64f81dca56e22f98e97a4c8e6c4f122bcdcb7e0d292cbfb18658c

Observation 1bca4ddf-b767-4fef-97ee-50634f531fc3 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.611434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.611434Z digest=sha256:3922d8cfa16f6dd230ad29f4b6b89610f8e83da1dd3a3fd873babd57968494c3

Observation 1c8bc8ab-dfff-4b38-b740-7d91927a1c8d · outbound

This paper cites Iterative Reasoning Preference Optimization.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Iterative Reasoning Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.682978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.682978Z digest=sha256:ee7fe23be8e21a2dcc102a35029493ce59976d89591b24e9302184bb4ecd63d1

Observation 0a63c277-c376-4034-8a15-bf4045abdc64 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.754574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.754574Z digest=sha256:cebbedda34e7e838a404cb64722a724036b28b66cfbf092446dc5560938c7fac

Observation 516444c0-147a-4c9c-94d0-d65198c0f411 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:41.665312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:38.829348Z digest=sha256:12202e7a2b4b7288a30837b12de2753c5b6c94737e2fc5909e805ee23638dab2

Observation f10f0bf1-7302-4892-b3a1-08e07374f506 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.904176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.904176Z digest=sha256:cb500638cb9efaeb9f73b2bee0f4d47f7e9b5df91566ba49c2a17f8aab91defc

Observation ebf9e83c-244e-42e8-87d9-acd9d69a2cdb · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:38.975908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:38.975908Z digest=sha256:c55641823783ac8dd2ff000fcad6b455a0bba42e30addf65338e7109e0063054

Observation 7dc98299-d56f-4f05-a29f-e51e7514e890 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.043292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.043292Z digest=sha256:c7637af2a65cbcb1af4e16161f84665dbf4716ea5f179d13a332a46a806a29e5

Observation 89a706e6-a3cf-4515-9476-3c59ddedfcdf · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:41.472975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:39.129896Z digest=sha256:94585db56025c22ef560510b555b501ffd38d2c73650d53391161443e4be8420

Observation ba407102-3452-4fc6-a313-f1dc0ebbda04 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Understanding the performance gap between online and offline alignment algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.184131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.184131Z digest=sha256:8e5ab070505a82e24bc85ad588f90bd7fef62dd8ea4d16357a3ae4377a31b48f

Observation 23e75e75-04d2-441f-87b8-ad46b8f76dcb · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.270943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.270943Z digest=sha256:e711ff09c66042f055331f53f6c88a03999793e307c5f648739ce0f98857150d

Observation 7cc21a04-c493-4bc8-b9c2-a7f276b5e96f · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Self-Play Preference Optimization for Language Model Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.353170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.353170Z digest=sha256:640dbb5ab2dda4867fab9b20bd03ae19320dbee6e12f836d9f4e6ae5b8497b1c

Observation e31260c0-e133-4221-9f08-b385be1d2d8f · outbound

This paper cites Minor DPO reject penalty to increase training robustness.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Minor DPO reject penalty to increase training robustness

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:52:40.355270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:39.447013Z digest=sha256:4c06dbafae44e860995db6dd2cc39ff548b57cf093b814a1f1a7c610d6894b76

Observation dd7e1b4f-6f3d-4a44-b87c-15b8f9510814 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:41.126372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:39.530754Z digest=sha256:b9109d5a88b5a72f51ebb3959d8d6d3dd45fb7b7d4659e8a8a1037b558780681

Observation b041ab72-18b7-4dc3-9737-9372bbad9760 · outbound

This paper cites Qwen2.5 Technical Report.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Qwen2.5 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.621098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.621098Z digest=sha256:14f0334faef49c155d8ee3cabfb1e1eb1bf6b37ca4097c765ec970261466c84b

Observation 1df16d6b-e59a-4f7a-a205-005bb957c8f0 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:40.912826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:39.681987Z digest=sha256:771c727127ccd626960219c8768026a2ad289be494810b00efed2031e18304bc

Observation 03dd5b3f-78d5-44f9-a972-9f66493b4e59 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.749756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.749756Z digest=sha256:57c8d049e826febfe502b34558157f5db0677e0f11962eaaf470669e0272d971

Observation 42593b94-453a-41e9-a4c5-655ca96d3b21 · outbound

This paper cites an unresolved cited work.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:52:40.726777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:52:39.889533Z digest=sha256:1e96f618070a8b461e0435aacecbde60861929149f9267327471dbf1dc2cc0fa

Observation b94d2f5b-1367-4867-abf6-632dd9e7d655 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment Instruction-Following Evaluation for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:39.972011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:39.972011Z digest=sha256:f1848fbe6265aaebc04659598d1b94d11d97bc5f0a199f64cef2d52575bef47e

Observation 3d3f3190-89d9-46ad-b0ad-c35eeb0ee1ef · outbound

This paper cites online" 'onlinestring :=.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment online" 'onlinestring :=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:40.073887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:40.073887Z digest=sha256:b2b573ced4f664b262ba7c88ecbc6ed6668c342ede8115d21fa6e88992e6f73e

Observation 63bc41f5-9f7f-4a0d-8fbd-12c65e1f724e · outbound

This paper cites write newline.

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:40.142142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:52:40.142142Z digest=sha256:8105b0ed06eb131e92aecef507243a7e7912d841f2948e2790ef40c52ca302b0

Pith citing papers

Observation b6c688fd-1aff-47c1-9b42-4fce800708bd · inbound

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs cites this paper.

Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs Rethinking DPO: The Role of Rejected Responses in Preference Misalignment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.079020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T09:52:20.508013Z digest=sha256:c93dd02d5eb99c413d2dd34ce51498a66fa2d11e0ad228aa6e60b4761f66408a

Observation 1b629639-8ed1-4636-8dfb-7d7e727e1127 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Rethinking DPO: The Role of Rejected Responses in Preference Misalignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:2584f949e35396c3c8c552113f3783a2f98e76c3fdfeae26543a14fb2df5807c