Pith. sign in

Paper Citation Record · LEDGER

BPO: Revisiting Preference Modeling in Direct Preference Optimization

As of 15 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.03557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03557 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:07:35.625982Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved14
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5950ee1-ad35-435a-b45e-7accb3e99048 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

BPO: Revisiting Preference Modeling in Direct Preference Optimization A general theoretical paradigm to understand learning from human preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.485605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.485605Z digest=sha256:772f6ea490c352c4c2f5e42d0a03baefda02e3e167303e72bd86da7cf6cca347

Observation f3974c4a-bf03-4ccf-9eb7-ad3e028f6862 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.490833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.490833Z digest=sha256:68fd6cb8f2ba432d3eb411b89196c93285139e6ba4b53c6e72a6100dfd781279

Observation e15453b3-d4dd-4897-8f13-766193e0a113 · outbound

This paper cites Bradley and Milton E.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Bradley and Milton E

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.133739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.496675Z digest=sha256:5141831e952c677166b3a2acda42221b6fb93d0b0403b2c76db442eca0fab046

Observation 8e2dfe17-bf77-4f5d-84bf-98628c0f0223 · outbound

This paper cites Christiano, Jan Leike, Tom B.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Christiano, Jan Leike, Tom B

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.117449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.501795Z digest=sha256:40aaf0e37b9886e637912ce81f75578f604474e7319a3e940ef2069cde2c5b2a

Observation 206d22a5-23d9-499f-845a-c3e075efd8a1 · outbound

This paper cites The Llama 3 Herd of Models.

BPO: Revisiting Preference Modeling in Direct Preference Optimization The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.511097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.511097Z digest=sha256:66d8f1f6ed260a7f01748f41f2ecaafaf2c5199ffb51e096234a609e93dd7127

Observation 7d895142-cc54-4c2f-9968-50909c9aae6e · outbound

This paper cites Model alignment as prospect theoretic optimization.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Model alignment as prospect theoretic optimization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.101078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.516121Z digest=sha256:dff8e2f0b1c098378eab01dc014480e0e457f654b19b8f3a3e05364c40091e84

Observation e5742135-7f18-4600-b5fa-a099ec757e78 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.084841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.521120Z digest=sha256:8321a2d55ca5f5c20ee4b912c42dca97a9b5016b6d01a1592eabd9db76bb31df

Observation 9d2fa7ae-b1fa-4363-862e-47de82b4b06a · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.526189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.526189Z digest=sha256:fda4d4b29e8c0800d6e6d2472c8664ddb2e80bc60d1d6c83b473c86bf312877b

Observation 0c5d7277-ac09-46ff-ba0b-e9ce2e6f52af · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.531217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.531217Z digest=sha256:b427a0ea082a40266571aa3588d3a70adb352a9c879d55377e2d740d9963bd09

Observation 9b54130d-a35a-43f2-8238-0200e750e39f · outbound

This paper cites GPT-4o System Card.

BPO: Revisiting Preference Modeling in Direct Preference Optimization GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.536884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.536884Z digest=sha256:ce5569d3f7891ece289e18619ee2fac55db06bd620996366754e77aa7327d0d8

Observation 42713f56-de29-4ef6-88d9-65a1bbef4e6b · outbound

This paper cites Smith, Yejin Choi, and Hanna Hajishirzi.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Smith, Yejin Choi, and Hanna Hajishirzi

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.057918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.541847Z digest=sha256:4df6de72f316d685acfd7b995bddcbbe397e6fc30a6fab4bc1b4848f8ded510d

Observation f48ca749-2555-4439-8d9b-b906c9031993 · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.040876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.546416Z digest=sha256:c9839e21da4ff3b8e6d4771497af785d497294b273c5bc5cbd2d5e02146b9495

Observation 473d34a4-8559-4511-b02f-13f0fb36e1f7 · outbound

This paper cites Mitigating the alignment tax of RLHF.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Mitigating the alignment tax of RLHF

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.017881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.550615Z digest=sha256:a237223aa64cd537420d88998e5667e6114903762cf6a073dd499407fcbbc19a

Observation d9fca5fb-4421-44c9-b46f-ce16871d4c3e · outbound

This paper cites Liu, and Jialu Liu.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Liu, and Jialu Liu

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.001899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.554964Z digest=sha256:f06401c396a79ea094a67eb8a20462172f37bde2facf70576c0cb0d5e30b8b67

Observation cd2b8c06-7a06-490c-919e-87236fc078f3 · outbound

This paper cites Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.563139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.563139Z digest=sha256:e223ead696065e97de40dc55847507e1025d20c436b60463d5874ea3c3686913

Observation 076167b2-c41a-4115-a719-d5d8db6f08a1 · outbound

This paper cites Mankowitz, Doina Precup, and Bilal Piot.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Mankowitz, Doina Precup, and Bilal Piot

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.970965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.567906Z digest=sha256:bc9e78fae01af9b3253aaac75ac0f71b2b8bc3c5366b21b2c47c29ddc67eabef

Observation 3600ca7a-a6cc-4e69-9115-59b0436bb9d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Training language models to follow instructions with human feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.956335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.572108Z digest=sha256:0170e57f445d1a85d0e4d9b0fcc1680d7db9ce08b4d821ec11cae2abd0ea4f54

Observation e2924f18-9bb9-4ea1-a250-7dba680a2c2a · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.576714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.576714Z digest=sha256:71632fe8b3613accff1fbb50298777d131f56f953c72e3d575b2c42c0ba4d96d

Observation f05344bb-5d1b-4041-ad1d-4162e15c0862 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Manning, Stefano Ermon, and Chelsea Finn

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.581612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.581612Z digest=sha256:af0f342b9eef50b9454924ee3feb776133fae19e789e6eaef5fe54b301f69a91

Observation 631117fe-0c0c-415f-9f38-53bf08418ae9 · outbound

This paper cites Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.586376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.586376Z digest=sha256:9e2b0b702f08a34b36d0a19eebf5f031fd901bad82a4d9af499623bc0edaa161

Observation 2e60a435-eee6-4934-9705-3076e36ab609 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.591108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.591108Z digest=sha256:3ce3ebf0d002d76128108bc3b177bd506174a6eb1d79a89d78ca8985da0fbef8

Observation f0838878-0a6b-45c3-a985-f6dd2b9d34f1 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Generalized preference optimization: A unified approach to offline alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.595929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.595929Z digest=sha256:50a4f49635b4c20a9f1c75350b01014903fe15024223914b0fc26651597c6aed

Observation cb73d625-993f-4164-9794-65d8dca20cb4 · outbound

This paper cites an unresolved cited work.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.600673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.600673Z digest=sha256:90507e07a852580beb58f630a8000361d7f593f4bf62a0f458b307a00bac1493

Observation df8e9244-c3c7-46ba-a09c-fdefdbd65d6f · outbound

This paper cites Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.889562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.605409Z digest=sha256:738429bf04c71e176188291e178e7112fcca8f0c66eee60f9bd032a99f966ace

Observation 30d500de-f698-43bd-a336-8e2b867414d7 · outbound

This paper cites Is DPO superior to PPO for LLM alignment? A comprehensive study.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Is DPO superior to PPO for LLM alignment? A comprehensive study

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.873245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.609927Z digest=sha256:fae07f45aefb53a4253f245e224df344b78e454f01ed289ebd425d432a6519c7

Observation ee9aeee2-2e93-44d4-981d-6d58fae6594e · outbound

This paper cites Qwen2.5 Technical Report.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.615063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.615063Z digest=sha256:5bcb578f021f2732ec64a6b6a0c26e58d333e65e5b7b0afde654db9570d2848e

Observation e928c41e-609e-4942-909e-de9dc75b6c28 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

BPO: Revisiting Preference Modeling in Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:07:35.619682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.619682Z digest=sha256:a9ed1b1ee581a18bc32f953ce4434cbe6b1495880551d1ad496b791078dc53e0

Observation 7aa043a6-3a0b-4006-8b71-1094c7bd69dd · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.855307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.625982Z digest=sha256:3c1565d1c1363c2f18ad1168cf19acfd386fed898d477a5674752dbfc41c18ba

Observation d9bb6598-2471-4399-a008-d5c62c8262ff · outbound

This paper cites an unresolved cited work.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T11:07:35.986800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:07:35.558965Z digest=sha256:ccab21dca40e06e8240280e60496eee3d106464c593321e8fde1497ad73eda21

Pith citing papers

No inbound Pith citation observations are available.