Pith. sign in

Paper Citation Record · LEDGER

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

As of 14 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 3 inbound Pith citation observations for arXiv:2412.14516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14516 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:16:32.070018Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.012360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:42:04.899091Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3043f17-7c1a-4907-9627-c1f5b1960ec7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.787908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.787908Z digest=sha256:94e2e4979acbc01386b7404233e5ba28f5b911d997fc6ed7b1350ef0ffca8766

Observation 7a53bf06-8323-4cfd-805a-41592a79ce75 · outbound

This paper cites Training language models to follow instructions with human feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.792312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.792312Z digest=sha256:57b0c42e725707b5362c72cadb23b28bca42a4441273dae42217730a087b07a2

Observation f1c5331e-62a6-4b14-8f7e-9d06a1242baf · outbound

This paper cites Learning to summarize with human feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to summarize with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.795485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.795485Z digest=sha256:5f7e7a1642a5c53724928d3e9e1b0ccaf86e4db4c9ab26415b25391a592f9d80

Observation c21db7c7-e901-4ace-aa9d-f03d596f7616 · outbound

This paper cites Deep reinforcement learning from human preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Deep reinforcement learning from human preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.798253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.798253Z digest=sha256:2ca6c3728eba49ab08fdc38fe73bdbe2bfba2f3b26663066cddb79a7e770e028

Observation 4ab6c605-1107-4c47-ba23-174ba61bb750 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.800935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.800935Z digest=sha256:56cc84c0b3f62fc393b4086f22256c5e256b6a57a3b41fd14d7142f7248cbb39

Observation c2e03700-8331-4b9a-9237-0832dace122f · outbound

This paper cites Implementation matters in deep policy gradients: A case study on ppo and trpo.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Implementation matters in deep policy gradients: A case study on ppo and trpo

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.805323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.805323Z digest=sha256:e3a0b1161ee2adb3b10b3bb500abd682e9f572ca873afb08fac7b86b525a8390

Observation c8364072-296c-4bfa-9f37-7b5b687dd2c0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.809910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.809910Z digest=sha256:2e15252ec888da4d7bde4d4f3b4b1f4d3926bc3b86ff87643f25d455efc25064

Observation de430066-de2d-42a1-aab7-fe0e94031552 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general theoretical paradigm to understand learning from human preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.813699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.813699Z digest=sha256:d9a4ae6eaab17919f6961e467e1783bddc50776ed9d5cfb52b50a5a3ed8f5fd4

Observation 252bde81-682b-48cd-b034-cb66fe3413f1 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.816513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.816513Z digest=sha256:6e889723e63fdeab1e88c7a69a1ba8e44988bb5610bea97b5506464c58a4f7ea

Observation 1260564c-4074-4593-9df5-ec7ae10168ed · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.820225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.820225Z digest=sha256:5af7b01f5c927fa92c8c0e93d2d0c0669f87505153603206cf954dde6e495cdf

Observation d69ef4d6-85ad-494f-a7e1-c9a4554f450f · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.823443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.823443Z digest=sha256:165e25034ca0b61145ddc9cf7b67a6804304f8e5c8855426ed4dec5332ba4c8c

Observation 99b3e0ae-f37d-4861-986f-cdd089439c71 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.827786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.827786Z digest=sha256:efe1a722c950893dde6453833245e0d714d839c13f1c746c64e80eba813d6358

Observation 6796dcc0-2ce7-4394-805f-483abf14982d · outbound

This paper cites Learning word vectors for sentiment analysis.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning word vectors for sentiment analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.831179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.831179Z digest=sha256:3c13b34a437f320554cf0fec4c1e82806f3c58feb3ce7706eae53154e1da690b

Observation 4cd60717-3a3a-459c-8c01-687001bbea5c · outbound

This paper cites Tl; dr: Mining reddit to learn automatic summarization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Tl; dr: Mining reddit to learn automatic summarization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.711917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.834698Z digest=sha256:c899e0b8b080eaf5a0bdfe8874ec82bc7f3386ea0acc6ab6206ace9e647f821f

Observation edadd08e-1c6a-43ed-998e-18509afea875 · outbound

This paper cites A framework for few-shot language model evaluation, 12 2023.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A framework for few-shot language model evaluation, 12 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.838747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.838747Z digest=sha256:7cbcd81f6c93359facd0ee2c6803a90670f827784464fa5c7a2f1e52ef904312

Observation 65c40c16-36fe-473b-ad35-f090bd37ad5b · outbound

This paper cites Nash Learning from Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Nash Learning from Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.841413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.841413Z digest=sha256:d9b1a645e52bbce8bf27b21e926acb475fe90ade299732f9cbc1cc405af146f7

Observation ec0e0994-acd2-4e42-a38b-9f8c176824b2 · outbound

This paper cites Statistical rejection sampling improves preference optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Statistical rejection sampling improves preference optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.700617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.844889Z digest=sha256:a9bd99203f0371332e4dad07ac43d9715685331d7ee2e2785d36e84017a881cf

Observation 70cce962-625c-40dd-ad7e-be59950a25ca · outbound

This paper cites Self-Rewarding Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Rewarding Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.847944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.847944Z digest=sha256:ffaa6b84500ff5f59b4f6b18d4b96f1d038c4352339675269c2d4f2490997cd9

Observation 0762ebe0-d00d-443c-ba4f-58f67230247e · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.851281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.851281Z digest=sha256:d51b7ed2008e68a93a9d43be09a095aaa4406e0d13f7716f90d829e2c8847629

Observation 6819dafb-6643-4018-b926-430f6136ac17 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.855210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.855210Z digest=sha256:b7838044d992db7bd987cd1ca1bc6c7b8116848fc1de0363cae8d79fa46a08b6

Observation 087d7327-30ac-4fe0-9c68-53e750be321b · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.859337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.859337Z digest=sha256:aa25bee19b57afa455d190f3e23a91ceb1fb6f67dfa94adf8e6c41b8050b7a84

Observation a7690f04-8146-4675-90d9-319bb5331e75 · outbound

This paper cites Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.693022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.863229Z digest=sha256:4cc1c842cc7b0b666bf8eb599057ac11bb9c1904f43021f9a6d818d679d80699

Observation 07476bde-f31a-476d-8daa-12c33e203891 · outbound

This paper cites Policy Optimization in RLHF: The Impact of Out-of-preference Data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Policy Optimization in RLHF: The Impact of Out-of-preference Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.865955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.865955Z digest=sha256:acfa96f959ea8902ad6ea21883cc941a8066ba2328e71595995bc62865e67ac1

Observation f3207bb2-2819-4a05-88bf-8338026d52c2 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.868865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.868865Z digest=sha256:db6a99c81f817acab354ae218200f6787fe527245005ffd7c03e4bede1b08010

Observation 00fe9418-3368-48a4-9310-8b5c3d35de57 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.872076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.872076Z digest=sha256:d4661e9111c53a6cd1a39b7ca4e419d4275cfe7cc068506e5abdb108a03649ce

Observation c89ca9c3-a17a-4814-871d-11c438522a3a · outbound

This paper cites Noise Contrastive Alignment of Language Models with Explicit Rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise Contrastive Alignment of Language Models with Explicit Rewards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.875936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.875936Z digest=sha256:dbc8ec82c8a383254cf8e08474df5cf6da9de7921a6846fc502a2a637d823c43

Observation 4960576a-471c-4581-801a-9b1107e24467 · outbound

This paper cites COPR: Continual Human Preference Learning via Optimal Policy Regularization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment COPR: Continual Human Preference Learning via Optimal Policy Regularization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:16:32.256071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.879139Z digest=sha256:d53708980852c163636aeeb78673eb44e6a1cbf16d578c954bdc64e7ab58c096

Observation c4044b74-173e-4ab9-8a56-2590f91da535 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Efficient Exact Optimization of Language Model Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.882201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.882201Z digest=sha256:555f4633cab9db023c6a5c84cf43e48bca056f23b384e81bed43a2bf647360a7

Observation 793ac3b3-1a83-4aea-9f02-c5230f473f25 · outbound

This paper cites Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.684069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.885458Z digest=sha256:3f5c1f4f4b9dfa09e0491439e3812ca38b9a8749ecf1fcf30991e991f433b049

Observation ea2b61fe-61ac-47a8-9026-a9ad5900318c · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.888937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.888937Z digest=sha256:0c189529f2a313a9ba41d0074953adb0fcb55310ee326f1fa8e01083330ff624

Observation 53babd5a-45e5-4f9d-8b08-949cb10d1c75 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.891745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.891745Z digest=sha256:7b835c6f18fada763a933a3be3583f3969b3fe05438bdc97325a57d4c1e4cda3

Observation f4659949-2b58-4b5b-8532-ade787f3a012 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.895081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.895081Z digest=sha256:c8751d62b48629bc7a8f637fae87a4382c83e85d5a01091240651401a7b5f2fa

Observation df554fa1-4c4e-4e47-8839-3d33f48fea7a · outbound

This paper cites Token-level Direct Preference Optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.898358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.898358Z digest=sha256:36e63f6a26340a67918e3ae7112de2258966803b9b3559e0cd6ed439533b83b4

Observation 1343b00a-f14a-4d03-ae9d-80f0692c6f1a · outbound

This paper cites A general offline reinforcement learning framework for interac- tive recommendation.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general offline reinforcement learning framework for interac- tive recommendation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.674683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.903506Z digest=sha256:94be9ff6e3b878580d7c69e215e46563acae8ec4be2c99684625a1f1a37084e0

Observation 6f6fb95b-bc85-4058-9a5f-7997b2df7c71 · outbound

This paper cites On calibration of modern neural networks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On calibration of modern neural networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.664634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.906673Z digest=sha256:2438acce3ef158b0827e4088664be80594cace3f393ff742faa9a2c86be641df

Observation 64781a79-dc7c-4d9e-82a9-962a674930b6 · outbound

This paper cites Scale calibration of deep ranking models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Scale calibration of deep ranking models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.655401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.910137Z digest=sha256:a05615485782dd84cf1c2497bbef7de095f8e7912b3813779403663b0bcbef52

Observation dfdcb30e-710b-4eb5-9e07-495b1f6bcc7d · outbound

This paper cites Calibrated model-based deep reinforcement learning.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Calibrated model-based deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.646066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.912682Z digest=sha256:dfab3e7b42c8d057f740821eb5a2c0d312b05c47cc78bcaa040d37d723199392

Observation 5714eafd-90fb-415e-b599-e941e60eb9b0 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.916373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.916373Z digest=sha256:f606446bb98acd2fa63a1a7e0af5384a3d53719593ae455b0ea99cb5f0744624

Observation 8437b2c8-1af5-4ddd-9ab1-d93fa3439987 · outbound

This paper cites On the Calibration of Large Language Models and Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On the Calibration of Large Language Models and Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.919947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.919947Z digest=sha256:6d8ec4eaa3f4b5c343b67388c2f8446611c40196310384a7f9d258b523214452

Observation 9312365d-b111-474d-b068-e16fa8b1bee6 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language Models (Mostly) Know What They Know

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.922990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.922990Z digest=sha256:2826a00bff271fffaa31439c9445e67d68c334ed9faadc8ee8eb9f54eeff915e

Observation 56d76232-f930-4dd0-af5b-52442956b335 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Rank analysis of incomplete block designs: I

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.631213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.927134Z digest=sha256:92344c33153782e8aebab45b1b8988f7f2e607fb69b27e7f6511b5e62aaf18ce

Observation 83463c92-b4e6-4ccd-8540-2d4bb110ed37 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.930130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.930130Z digest=sha256:28ecd813693b43fc978adbabf9373fe76ce1a2b3eb7225947c138be075186942

Observation a4256de8-453d-45da-95c8-e554d4f05985 · outbound

This paper cites How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.933188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.933188Z digest=sha256:9c81e33597633ad7c6488f51813cfdc5f92ad654b415deb59f314e856cfce292

Observation 0943c1fb-e0db-48ed-9cf5-9f167b4771ef · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.937213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.937213Z digest=sha256:633c390546f0aec9a5e15c3dcd0fb8cefb31740bd2c3982e3adb1b61b2ca5c41

Observation ac1a2811-daad-4d69-bb65-a0193be5fac8 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reinforced Self-Training (ReST) for Language Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.940300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.940300Z digest=sha256:38543aadc657d573736726395260e0ac9aa24535e9fad664a49a6b860388f7bf

Observation 25e9db84-767a-48ea-b70d-e03cd68fe5a9 · outbound

This paper cites Openchat: Advancing open-source language models with mixed-quality data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Openchat: Advancing open-source language models with mixed-quality data

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.621160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.944005Z digest=sha256:3c9cf359f6c8a2d7299aece7fa70221fe60d44ed2d59e30a6936888ffc4bc465

Observation 84b5bfc4-f444-45dd-ac05-f508bab73996 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.947063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.947063Z digest=sha256:b42d873f00fbfda4ffb99cfb3dc0bc99a7c9fc73ccbd061aaadaee6ec3747f6a

Observation fa582487-b8bf-4e08-9c55-7a73db68d3b8 · outbound

This paper cites Machine learning: a probabilistic perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Machine learning: a probabilistic perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.949861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.949861Z digest=sha256:009ce0b68802bc7297049a399fcbe97eedf0bfddcba920e55398c594401e48cd

Observation 42d83879-5ccc-4995-84eb-62af187d97d4 · outbound

This paper cites Improving Policy Gradient by Exploring Under-appreciated Rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Improving Policy Gradient by Exploring Under-appreciated Rewards

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.952713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.952713Z digest=sha256:64890aa83c0ab84415c1f56b1a6e065746c57c466be5e151f0e5705ab09f487e

Observation 2a03ab81-e9a4-4ed8-a15e-156dba954253 · outbound

This paper cites Learning how to propagate messages in graph neural networks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning how to propagate messages in graph neural networks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.602814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.956300Z digest=sha256:86fa5b0782d23a7e998291d01294ec6b4f4d2127b14a075f3bce10d7ab53b443

Observation 00ecfb45-e50e-47c1-b8fa-c9e7af06eb87 · outbound

This paper cites Learning to generalize from sparse and underspecified rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to generalize from sparse and underspecified rewards

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.591645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.958843Z digest=sha256:74b861b09340966ef08b1ae18584ce0a8410b872c55272ccc06c6b52fa85512e

Observation f71633dc-4925-4aad-8ec1-3b59b4422e85 · outbound

This paper cites Decoupled self-supervised learning for graphs.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Decoupled self-supervised learning for graphs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.581460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.961637Z digest=sha256:6c82f39c8ca8e4e192a8ffea535c1d2eea2987cf88205e639fe4815acf60f543

Observation f57ea6f7-d3af-4c71-b8aa-df3d3e2ac0ef · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.964333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.964333Z digest=sha256:ed52aec3fcbfe24298c36df3b53f2296f348db771367b73b79b55a685794afc9

Observation 5a888593-3615-4760-a93b-621af2650ef9 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Zephyr: Direct Distillation of LM Alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.967147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.967147Z digest=sha256:0405aa091a21dc5bca23762ea76dee2dc2d4c6e7c021d29b6d1ad05ec3aaacf3

Observation 8842ab19-1cfe-4140-910d-d92946430a09 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.970515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.970515Z digest=sha256:a0da2fbd52ca6187e4554deefe512d756c47ec4873b8ad53142d58d47937be2b

Observation 8a528a66-0fca-47d9-91f4-689ff083e0ae · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.973598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.973598Z digest=sha256:4fba998b889ede45ad6e21657dedb705b9364f8bd029387112bf9cd46b64adbf

Observation 53e1e8af-cde6-4422-ac5d-4e030c0db960 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Instruction-Following Evaluation for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.977059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.977059Z digest=sha256:e356c7c0ac2019b89e76198d81e0a704dc32cd3266801bdd8c2b9b8963a8d9f3

Observation 38d0ece0-5b9d-4acd-89ef-12140dc30d40 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.980152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.980152Z digest=sha256:792a6792cf642ca63204caae96824af5820a3db2ea43b5887a92ac72f4ce7153

Observation 9f242457-b2f7-4380-9e7b-3bcb877d3602 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.983210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.983210Z digest=sha256:33294ed5a3faa55f59b33857085c66702ec97be4c6ee350245bff26172648e38

Observation 085c13ee-d299-423b-b2c1-408a20437e10 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training Verifiers to Solve Math Word Problems

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.986326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.986326Z digest=sha256:179c325feda37ffa9fa45558950204ab0b4faa7faa0b255e33614a809f2a58d3

Observation f68d5328-0d1d-41ed-86d2-480ad6135bdf · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Measuring Mathematical Problem Solving With the MATH Dataset

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.989194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.989194Z digest=sha256:d31aa1f4467b4ceb8f69baf88911a1669d26d8b5a664f15df3325be0a6204be5

Observation 959e2286-93db-4f1d-bf5b-cc7288c8181c · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.991750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.991750Z digest=sha256:7a8681e437b7e7d7e29d79ad12ff7eeab938bc400e06a59e79a1d2d8f398bba0

Observation b5c75b4a-1225-4dd4-bbf8-58a1b9a2280f · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Pythia: A suite for analyzing large language models across training and scaling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.994540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.994540Z digest=sha256:3ce90d70e66edf5dae8714eefc4cabd982712d09f46e0b6ebc7a404fcd800d4c

Observation 3eecb998-3173-402e-b8d0-ea74583ca919 · outbound

This paper cites Language models are unsupervised multitask learners.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language models are unsupervised multitask learners

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.559450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:31.997113Z digest=sha256:193bb0d85bd47cb801935a441abe8938c233de4aaa252b4a619d100365ea0a8b

Observation d2f9f718-aa82-40fd-96c9-61af13d31330 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.999493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.999493Z digest=sha256:36cdb78d962f29e4d67afacd3c1588cc84b766160cfa411eb4c592e1d80caaee

Observation dc65e0f0-fbbf-4f39-a91c-6616e4399414 · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.549991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.003524Z digest=sha256:2a5c8c8f33405af877e80b5dbc3ae647b165e9eed1c2946360d58e8b222c2b4d

Observation 650830fe-ea77-4a4c-b4dd-a0dc74095e6d · outbound

This paper cites Iterative Reasoning Preference Optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Reasoning Preference Optimization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.006296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.006296Z digest=sha256:2b94842d37da5fa068533b41a04c2286b4ce1c5011a94f1906ada269026938c3

Observation 33bd544e-d5b4-4168-b0a0-581e3d3d6058 · outbound

This paper cites Information, divergence and risk for binary experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Information, divergence and risk for binary experiments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.540757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.010038Z digest=sha256:95c7e0685bdb7816c274635eec6722df16b27cd489b02aaceb9d775bbecf3786

Observation da37e0fa-1e0f-437f-9935-12465a4d6df0 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reward augmented maximum likelihood for neural structured prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.530302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.012829Z digest=sha256:2676865473d2d5e5c473f8e8023b5173d0e6ad73278e1d0e3f04cdb9040f5058

Observation 05b24dda-9261-44f0-9a92-abeb00905921 · outbound

This paper cites Clustering with bregman divergences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Clustering with bregman divergences

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.520990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.016829Z digest=sha256:9c7754a8e5af536443fed4ecb8f64fcf32d23d9c8a206d7ff09b0238dafed024

Observation 41168029-a0b7-4045-82c1-31903ff9efa5 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.019439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.019439Z digest=sha256:59f95d8ea02720dcfe16c2c6aa59ec225f5ab6ca976c22026ed3ae09f5184b20

Observation 484e4191-a869-4b78-bbf9-af172e8cac7c · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.022520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.022520Z digest=sha256:cbc1f556f50afde5fb33641d5a40b3dd89f01bed956aa2e4338fa6c0b7db665a

Observation 8210680e-89c7-418e-bff8-ce1a1fcc093d · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.511971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.027443Z digest=sha256:2b1fc4e136b8cfe6418c66b2a4403d1e4e0812820e6b36e2d2515eff7f6fb5ba

Observation 26608078-d248-42dd-abf9-23a4eb8d7b11 · outbound

This paper cites Limitations.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Limitations

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.501635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.030919Z digest=sha256:088206afdbbb7132fbec5ec0304b04aebd6dbb376811f62d071348cfddba45c8

Observation 4884cd23-33a9-4595-8b34-3bae026efaf8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.489123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.034227Z digest=sha256:95d1393a022ac4f2c4b2af2f55ca0166af2ccbc59fa9d8867024e1af70029041

Observation 4826bd82-96f4-4751-875e-7cee43bda8b5 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.476708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.038330Z digest=sha256:7d6d263ddbcb95c12fa1cbfdaedddba9cf41ab40a540482f467c9758cd6624e1

Observation c0db1cac-3a60-4de3-8d6a-1daad2ad7748 · outbound

This paper cites • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.467290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.041960Z digest=sha256:2e414833f4cf08c200eea14f3197504978d369650e4d8a1ca97dd8de339625e1

Observation aec86a50-2589-4464-aa65-e396144e2d65 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.455381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.044849Z digest=sha256:79b6a96b37eceb6db8a487442cdb5b64294616c71d807356fa0f8de98f4f570d

Observation 06642fd1-12b3-4dcd-bb92-f67df389f24a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.446467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.047698Z digest=sha256:e19a75b06e973fb530c1ee26e5d8aabb2e34a2fc12bb06f2f81238524d578246

Observation 28738ff5-4694-48a8-8d58-87fd63ab9128 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.437364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.050305Z digest=sha256:2cf18fee4b5c7cef6ab6505b315c7cb912c11e3efa7408719c3bf3576714c46b

Observation 30282f20-6d4c-4f28-aed4-c4af1e1fea3c · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.426109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.052865Z digest=sha256:df654f59683f104b231e91631bc9e7ea72c306441a266412b72c60dcb4f3fe1c

Observation f9cea911-8a11-41e2-917e-369700efd3ee · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.414192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.055941Z digest=sha256:1a06e4f00e6adade9a05d8fb00e29537069c48dbc800b58b036c3a9a101f1117

Observation 9bb2e117-02d4-49b7-b2ec-819966320e54 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper poses no such risks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.058383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.058383Z digest=sha256:e289485634943312dd681442887c003b24040007e340287070cd862b5dde6f66

Observation ea19d25d-37db-4e40-9eea-312ff696fdf3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not use existing assets

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.397363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.060850Z digest=sha256:eff65c8690cd01d737fcc23a7c268a3dee90e3453fcaebae0520324d590c4867

Observation 76ebee73-8f0e-4919-9cc8-415d8e5c5bb6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not release new assets

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.064064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.064064Z digest=sha256:aac5a086211cfc4b33cbf545853aa2a6f809f6750dfcec57b1d9c91036609021

Observation 5df0804e-650e-45aa-91b6-34563bce2854 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.067064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.067064Z digest=sha256:f7e744b43c49f1c627cc9646bd7672e55961cba0b12d69dfb07c7f5bc6234682

Observation ea1687f3-4a6e-4d5d-aeef-083d0bd918a7 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.375743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T12:16:32.070018Z digest=sha256:59b3894ddba082abd3ec0d95ac33751d0d7a6d0bbb2a9a87babf27d8799b9678

Pith citing papers

Observation 5983e2f6-f1c3-49aa-999c-9598cfd3d96f · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.012360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.012360Z digest=sha256:b54f2024673d40328d7bfc986ce0f14d6c83ecf4ff3704b99f778bad687e88cb

Observation 4b664141-2b19-4be9-bc82-8fb1de2a8e68 · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.901351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:f1074c21f4b51e1f3ed5e4be7220b2533574d248497d7c599bfffae3e0f6b2dc

Observation 10e735fa-5805-42b3-a9da-57eff2c8f526 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.691864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:2328690398a370618fef0c329a824146eb6f4986aa3660b70d64b93eb39a354d