Pith. sign in

Paper Citation Record · LEDGER

DPO-Shift: Shifting the Distribution of Direct Preference Optimization

As of 24 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2502.07599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07599 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.052839Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.454892Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:28:54.031274Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33145946-9ccf-4395-941c-5093e741f713 · outbound

This paper cites GPT-4 Technical Report.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.840955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.840955Z digest=sha256:926d5f12d27551a51b39c638f598ef15739c084704f972e79c9cc2ec322a8326

Observation 63faaea3-55da-418b-a837-b8b4b46261b3 · outbound

This paper cites Llama 3 model card.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Llama 3 model card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.846902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.846902Z digest=sha256:b999b8ecc82ee28eeb1850ae25ef558d81fb38ac94a3772f6eefbee880112c64

Observation 6fa528fe-52bd-42ff-9cb5-4580d10ffaa2 · outbound

This paper cites Capybara-preferences dataset card.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Capybara-preferences dataset card

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.779862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.851997Z digest=sha256:e8477637dddacaeb0aa4e6ad9d94448fbba6827003cc846020095850fc2dcba0

Observation ba312bda-118d-480d-a798-75fbcb774058 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization A general theoretical paradigm to understand learning from human preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.856797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.856797Z digest=sha256:12216cc3fb7693a204f14a424483993feae55e6caf237b594018f3bbee8ce698

Observation 5b789d96-2093-4d08-9c54-ec647ad0fa42 · outbound

This paper cites Qwen Technical Report.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.861587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.861587Z digest=sha256:dad00687c88400293de55eb6f749e54e75ade08a1034726ea4417a943e517110

Observation a1c0063c-2605-4660-85a2-e3cd44243ef6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.867014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.867014Z digest=sha256:7474322ff77d6243c8f34dc674670cc789bf4344cddb6dc0017a945c499659da

Observation 34ef2061-aab6-4200-b693-268954668416 · outbound

This paper cites an unresolved cited work.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.873041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.873041Z digest=sha256:bb7dbf198986b233aa5e79e0699507c8ecfca3562cdd340683d1cde17533f6e7

Observation 148100a8-e34a-4912-b8ec-1dc0f6cf402f · outbound

This paper cites Evaluation metrics for language models.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Evaluation metrics for language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.745304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.878236Z digest=sha256:3b552a658eccf4ac0457a406f3824dc0917002bcff997adf2349607f4cd5a8b6

Observation 48da12f1-a16a-4479-a2c2-b52e41046536 · outbound

This paper cites Deep reinforcement learning from human preferences.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.882971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.882971Z digest=sha256:7006b4bbc795bafe5bb0c6638aa515e5f0619f9bc093687f29abb09e72417ee3

Observation 02f4ec9a-402a-493b-9787-10a3d55970ce · outbound

This paper cites Perplexed: Understanding When Large Language Models are Confused.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Perplexed: Understanding When Large Language Models are Confused

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:17:46.496898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.888313Z digest=sha256:73b463cfbc2e0619bea9e2bf785a5aa89592117940887b2b27b6133be8ec6db9

Observation 10a41111-4a58-4256-869c-4013fd5d9ca8 · outbound

This paper cites UltraFeedback: Boosting language models with high-quality feedback.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization UltraFeedback: Boosting language models with high-quality feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.720553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.893619Z digest=sha256:46d365e0055ef4ffdd3a3b7073faed595a454511cb8a8ad92c2d50be07fc21ec

Observation bbbc76a2-5801-4144-b6a0-5db54b7d0604 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Enhancing chat language models by scaling high-quality instructional conversations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.704877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.897907Z digest=sha256:91f53aebd3b91161377dc0b2267a2bb41b4ad1a69fe41cbf72f5b35664f261e3

Observation 2396ef46-2d0e-4e60-8dbe-24000f16334e · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.907909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.907909Z digest=sha256:52bb12d87bdd45a67def4612a22a3045ff48eb9ff5a46b28db7b084a14b2a46c

Observation 39c339e4-edd7-42fb-97ce-a3791d515913 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization ORPO: Monolithic Preference Optimization without Reference Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.912780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.912780Z digest=sha256:bc550021810735f5f7662ffef4a232ebb2a13e29baefc77a547a01ac0b89eda9

Observation 76184030-189c-46ef-918c-4bfd064b0068 · outbound

This paper cites SDPO: Segment-Level Direct Preference Optimization for Social Agents.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization SDPO: Segment-Level Direct Preference Optimization for Social Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.918129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.918129Z digest=sha256:5f3acf153197e9821c90210a7410a92d637476f6b06179c44c2ecdaf40e1e54f

Observation b136f837-6d43-4c34-937a-3bc5eb491de4 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.923786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.923786Z digest=sha256:9488b5692629344fc9526ac2e354c861300407d1ebf08a3a2e4e8090f9892d53

Observation 5aa3f7ec-4432-493e-98b6-60b849829c99 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.928601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.928601Z digest=sha256:294d2367a7cfab2b0521a04042f39ba891326c8df0cd599b38b1e618864789d3

Observation 7c695b3f-a832-4f63-a452-1d08566350ee · outbound

This paper cites Entropy Controllable Direct Preference Optimization.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Entropy Controllable Direct Preference Optimization

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:17:46.400700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.934605Z digest=sha256:6d1d1ac7ab669d777666f4cba2fc80a2ff98998f580048efa76d7f1034ccb627

Observation c14b757c-8579-4c70-af67-3c2f211be6be · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.689536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:45.939314Z digest=sha256:b3b487ae99324975cc281c81151db5943ac3a251b07c19e1a8c04017f2453f67

Observation f0d12657-a7e6-4a86-b414-7aeefe75b728 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.943180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.943180Z digest=sha256:5ffb2390b850cb6b8130adcb29547efe6e0674deab54aa09d2477507dca6d75f

Observation a9b9d3f8-c4db-4d75-a847-131dd26a31a9 · outbound

This paper cites Pre-dpo: Improving data utilization in direct preference optimization using a guiding reference model.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Pre-dpo: Improving data utilization in direct preference optimization using a guiding reference model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.949210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.949210Z digest=sha256:e65598f02873a432a863308fd650d504161bc6592d65bae565e44bfd73de0a99

Observation 65119daf-0aba-4cb7-a45a-6001e6a64167 · outbound

This paper cites Iterative Reasoning Preference Optimization.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Iterative Reasoning Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.953811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.953811Z digest=sha256:d9dd41dec3c1731c6e5bc9b50b74ea3881d71af2e0ed388293c864e36e813693

Observation 5cc0fbaa-b4f8-4303-a375-4b58990c6ebf · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.959380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.959380Z digest=sha256:97b9a978db29356a10b19a8c8e9076168c6fdb06e11f48d027fa84a28c54e460

Observation 3a6b671e-0090-4735-9254-5d340e7d76f7 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.964514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.964514Z digest=sha256:e95a9270540bfb5c19be4e8b613407aafda21c288e0de1d68ec553dacc3ebf4b

Observation fcdb08b6-9900-4f8e-a1a3-43257109423c · outbound

This paper cites Manning, and Chelsea Finn.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Manning, and Chelsea Finn

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.969176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.969176Z digest=sha256:8efaa44226426bc18b3fae225c61d202aa78054a86dde351a4dc6145f2b0e2a7

Observation d63e46b8-c36d-462a-b4de-750bbc12eb94 · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.973947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.973947Z digest=sha256:c93618b09191d890b3947f28a4cf4a9e78247eec189a36299d4e1e93417899b8

Observation 7acb93ba-52eb-43f3-aea2-2b3d17db54a2 · outbound

This paper cites Learning to summarize with human feedback.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Learning to summarize with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.979808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.979808Z digest=sha256:fd6b4402b1d2abf58b926c000ea06dc6579c8cc4ca6c1d2bb3e8a1bb9fd60a5f

Observation 26f9010d-7322-4a94-9531-3db1be88ae8c · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.985439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.985439Z digest=sha256:448dab9533d27344a7f7eab76db19a6eb561ec692c5dcca4e3de4d3c55634e55

Observation 1708aeca-820c-4d29-93d4-736cf1eb84b1 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.991858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.991858Z digest=sha256:32143182e80ca8c2e130aa39549d231b69231f2d01b3301d3f3eaaef0cc6c822

Observation 13350852-ce03-4237-80ce-28b34add8a2a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.998038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.998038Z digest=sha256:a1375f0e14325611c7129d9c51e7a84cbb715b28cbeed5c7324b5908fd48eb7b

Observation 7448a94a-bc89-4a02-a58b-c3832af6a8ef · outbound

This paper cites Rush, and Thomas Wolf.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Rush, and Thomas Wolf

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.640454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.003075Z digest=sha256:83d1ae0b7df16e0fed6e0d7f8e27d38a2104b37e02f5677d00cd8843abda0419

Observation 4a88f863-0b9b-4fe8-9650-f4f08dd9e28e · outbound

This paper cites SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:17:46.162294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.007691Z digest=sha256:11337dd8e6cff223db70bd1fe6e99ce1420de51d98116c6ecf50d22f116de3bb

Observation 5983e2f6-f1c3-49aa-999c-9598cfd3d96f · outbound

This paper cites Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.012360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.012360Z digest=sha256:ddc3b09065b5dde49b25737e27cd2a716302ae03b7d16d713af7a858d52f5e38

Observation 631602a3-3634-47d5-a68d-f8891261281f · outbound

This paper cites A systematic evaluation of large language models of code.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization A systematic evaluation of large language models of code

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.621940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.018017Z digest=sha256:333cb2291433afc9cad51dcc73d5292c0a46616cd9d83006347f52df6a088366

Observation 6b21bd03-5d6a-434e-8f60-6a9969484189 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.022685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.022685Z digest=sha256:46af47fcdae6eb55b02460040e2f8d8617915206761d224a26ee84ab206955c9

Observation d3486637-e626-491e-b325-76d1a306fcd6 · outbound

This paper cites Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.028385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.028385Z digest=sha256:ac5458c84b63e9e1b46a890bf711af0fc4c13536e32e19b4061a048eb1163f39

Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.033521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.033521Z digest=sha256:b2538218ae59b649dab2812a892cc8bc6f9f36955014bca571d874735b061c55

Observation 5b6cb968-5e4e-46a8-b591-1bf13cb00973 · outbound

This paper cites Judging LLM-as-a-judge with MT-Bench and Chatbot Arena.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Judging LLM-as-a-judge with MT-Bench and Chatbot Arena

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:17:46.606898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.038408Z digest=sha256:32f097503396ea37ba0fbac28a0b5a48c8f7fa41a62b15dadb2253683b61b773

Observation b26182b0-e13d-47b2-8538-0bb5063d22d7 · outbound

This paper cites an unresolved cited work.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:17:46.591538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.043297Z digest=sha256:fc16e70d4e9f9b59c436cf8539ff3baa6b720d261523cc56607fc14b8bea943e

Observation b91271ad-6335-44da-a823-efee30f9354a · outbound

This paper cites an unresolved cited work.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:17:46.576011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.048675Z digest=sha256:d81fa9bf77ab5da39c42ccdb989c5e643b8172e4c9a87e1e97f0ce1fdf72d34b

Observation 4e3f4637-e232-4ed4-b1e0-d0973781cebb · outbound

This paper cites an unresolved cited work.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:17:46.559467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T12:17:46.052839Z digest=sha256:f45e4395a66187fcc66f4eb2ced8cf3d1d56e49da6455508f3c4394bfa593d6e

Pith citing papers

Observation 5d220a40-9759-4e58-a761-1d70cbc63392 · inbound

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment cites this paper.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment DPO-Shift: Shifting the Distribution of Direct Preference Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.454892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.454892Z digest=sha256:a5a403266efe85cedbb183918bdab3aeae7921389e42419b4a856d506fa086c5

Observation 9ec5bba3-48a6-487e-b041-1bc3367e23d9 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs DPO-Shift: Shifting the Distribution of Direct Preference Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:28:54.149778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T10:28:47.271156Z digest=sha256:5d6e7c14cfc51dd54d68df7edb13a5b70474d1a63836f57893e0cecf7e9c15f5