Pith. sign in

Paper Citation Record · LEDGER

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable)

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2507.07855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07855 v4

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:38.826438Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 965346be-dd5c-4541-a403-540a244512fa · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.129033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.515815Z digest=sha256:834d76429ead012047cef2f341fde2b5f18f0b49a0ef1250c92abf7925704d12

Observation 81fccb27-e677-4950-bdec-22a72fff3fde · outbound

This paper cites Alfano, S.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Alfano, S

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:52:39.375745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.577577Z digest=sha256:6fa23d94da51e998abdb8183ebcbb2dc9d0f43ed571bb691fccf3fe1e6b506de

Observation 80ac216e-2012-4d87-b9ef-03a0217c83c2 · outbound

This paper cites Amari and H.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Amari and H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.116234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.693550Z digest=sha256:4ebd2e3c30bf728f5a6222e01cf4dcc842552eedfae205d7f5275d79beceab11

Observation 9a063b77-0699-426f-9c4c-ff02860b399c · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.103357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.787534Z digest=sha256:61700d95db3e69fa57e280fd4d0d64c0a3ed69e89f05725ec2cb6b70c8f049ba

Observation d8a5099f-8521-45ab-8fd5-6f6d36949e01 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.091265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.863589Z digest=sha256:43d8abb399b4d819d2fb8ce9c2e83dd2a8d0e1a21ed3576c5b4c4cc45f671801

Observation 5acc9787-cfeb-4a03-a9ef-73a41d7caafc · outbound

This paper cites Bao and N.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Bao and N

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.078926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.934402Z digest=sha256:4350e1de5fd7c931da9e314ce6449a0a535aebefbc6932c01b35417a77c8e3f6

Observation 75ca5948-f075-4972-aad3-1adc2626e9d7 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.065896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.039406Z digest=sha256:218afb83a7df842f2cbee4ed18c38820f366cc5c9e513bfc97d036cd01793eea

Observation 02724bf2-6657-422a-991a-34a97ae0e295 · outbound

This paper cites Blondel, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Blondel, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.053423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.197979Z digest=sha256:04a0f7022c9a1c2a5d2065b6a75d6dc68387e1e698710eec00d3dc97d744c375

Observation 9f798169-da66-4976-988e-1f5427eabfa9 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.040610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.263682Z digest=sha256:6979f223a0c08d0bc89a492ba81b571500320dfc5499d94152350ff483254147

Observation 02b8a3f9-e2c3-4d3b-9b5f-1ec9f1231fd8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.027737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.366115Z digest=sha256:71417e1f1fc4501bdc3b3910191837a648f3f3b48fe8c2eb8a22da299ec97e6e

Observation 98eb9345-e399-4045-bbdb-d7f2cdf0b977 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.014204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.469748Z digest=sha256:40de9762973a9391fc6c04f2e88d703b5c8790af65ae7334f302f156bd82693b

Observation ec68244f-d308-4ea8-8d49-23025093a4f6 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.999814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.542796Z digest=sha256:e1c84dceb271359caf32e70201773b642f0c62e29c90d142d38606bed8e06f28

Observation c8a8fa34-7bb3-4ca3-9565-c24ec8214238 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.985297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.639454Z digest=sha256:11f43867ae7c4c6dfa407ede736c6762e4186fdb5772e5ae428e3b0029e40c37

Observation d5c96902-2571-47f1-997d-aba9d37f84e3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.971992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.793576Z digest=sha256:4d1ea3848863bdefe410a2e830a037d8924ae272fd8c154f94be91d2e4222182

Observation 3aba22c1-3715-4e1a-8acc-44cf78ad5ff3 · outbound

This paper cites Doignon and J.-C.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Doignon and J.-C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.958070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.870271Z digest=sha256:476df9893207a4c1eaf222d61b191883dee29d571824f48088767d82993b63eb

Observation e6e2ce3d-05c8-46d8-99fd-a7221203e38f · outbound

This paper cites Ethayarajh, W.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Ethayarajh, W

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.945039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.946745Z digest=sha256:74fb2e85d42a05f0871628da2e674b13d6dada48a636c7c665c9a9fbc15b0f85

Observation b4cd4805-d636-4f9b-999f-2bf556a9bfbe · outbound

This paper cites Gneiting and A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Gneiting and A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.931857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.057746Z digest=sha256:7fb25f5c00960e51430b8f400167a2a52d6587a184670aed80da86a58a4b92f6

Observation 32cf51ff-128e-4429-903c-9e18fb93f4dc · outbound

This paper cites The Llama 3 Herd of Models.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.156728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.156728Z digest=sha256:95bd703c475421104541a2e363cb3e7cc583102b0bb0cd98e0f8955ce4780861

Observation 0b4a9192-b9bf-4986-bc94-d7c986731020 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.231062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.231062Z digest=sha256:bbe4f6a21ae70597cbc34937e2d01fa3bf801782122c4d8bf2481b1de884b470

Observation 8a966e47-8557-4927-a89b-76ffe185a115 · outbound

This paper cites AlphaPO: Reward Shape Matters for LLM Alignment.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) AlphaPO: Reward Shape Matters for LLM Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.347941Z digest=sha256:6dc531102f4334140ed9595f95ad4eab3370f3cba11bdaf5ead4668ff46a6f1c

Observation 0b463851-738e-4987-a460-caae629ac756 · outbound

This paper cites Hastie, R.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Hastie, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.909162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.450714Z digest=sha256:e3d8af9b795e7a722ac478e9b0d45ddf17820eaf00ce2d74f9abe1364346dd96

Observation 5229e8a3-3a7e-4622-99b7-61e3fb4701fb · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.896574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.600101Z digest=sha256:4037fc8671463db60af14eeb5ead437d79785e96a1cf408edcba473dfccbaae8

Observation 8a5b238e-5927-42cc-bb41-8ba86e955dda · outbound

This paper cites Huang, W.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Huang, W

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.883467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.670121Z digest=sha256:cdb860fc0c08cf207af56b5e5ca0097c0f509ac87506de3f44bc0e341555d0c1

Observation f4ce977c-d379-4f66-ae69-9b6c4f9dd52b · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.793726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.793726Z digest=sha256:b6b6a35f9468584f56f9ce0ccd2f2bbfb926d0f5396a10cd93d73cd1c52fe227

Observation 54ec5f14-ff72-42b4-be43-e7ced6156b14 · outbound

This paper cites Kakade, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Kakade, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.870296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.884856Z digest=sha256:ba1eb4bafc43586a973dd8410ff8c2e0c0ea02557b72bc9fd235a789947cf8d9

Observation 207c34d9-a568-4712-a475-d5c6144f41cf · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.857109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.950676Z digest=sha256:3a2e78753ad12355a842808e0bd5daf61484052180cd9e96ad1d3f2fcc2e1d49

Observation 1a545f81-5c72-4731-a87e-71a3aea4a00b · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.843801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.013640Z digest=sha256:4da9630f7eae1778b03c24d318c2f4844221a69e7d748462d8179f61f4a72db0

Observation d2c9642c-e399-49e0-952e-8cf8db96b42a · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.830449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.081485Z digest=sha256:d58dee5372151be7011156fcbfaeac18a78ee266d9b156a455fb2c6b8750f698

Observation d31023e5-7897-444b-a6ca-3cbc8878500c · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Direct Preference Knowledge Distillation for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.169261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:36.169261Z digest=sha256:ceab59ed9402672579800cbbddd4f8dadf2390c64a38dcb72ad39b4b8b6cd5cf

Observation 90904036-3fb3-48b8-8851-53fe7c53e6a8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.817390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.262670Z digest=sha256:f06a6dd125ad3711ad275ff0cce9c61483f83c31682b05ebd863bda14c5e9b45

Observation 8a726a5b-64ff-4d3e-b2b1-3d1b9561d547 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.804388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.334781Z digest=sha256:741a61bfda5bfc674d71c693e73b93a274d1d8b71bf9a62e4ed9410b8ea5000f

Observation 84697ba7-17f3-401a-adc4-40948e8236d8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.791682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.408635Z digest=sha256:fcaf675e8912fdea1a700605376e80947be6b6d358ba84b6a7a26448a5d0b3f1

Observation 8ba6d603-4793-466b-9b3d-f8b415ccc5aa · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.778771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.471253Z digest=sha256:e2fe61b17a156255881bfe37e8001768fab93ba42a95d2f4f1437f02979cd586

Observation 076ddc84-144c-4949-81a0-2359698b150f · outbound

This paper cites McCarthy.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) McCarthy

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.765690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.561030Z digest=sha256:cdc45a3b42fa9ec2012ade7eb85856c437cc42d10f557177c227da31ef11992d

Observation 30c1683d-dab5-4f38-9829-ae7326836ebf · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.752867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.626119Z digest=sha256:bb498f031fed023c6481054151589257d756dde4e52f0a96575a8e768a44206e

Observation 2d8a6005-dce8-457c-94ce-b0e00586c136 · outbound

This paper cites Mitchell.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Mitchell

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.740328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.721352Z digest=sha256:f2fce79e97c308e86758502a03ac0f79fa26d85b1ec20d6d12ddb1e57fa3b01e

Observation 640d0667-d9b3-4cf8-a373-16a7155fd115 · outbound

This paper cites Nock and A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.726540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.788286Z digest=sha256:f91fb7ce5ea1313f164c72d260c94a7efda5ed1d7a8bf11a63e738d8cf96c86d

Observation 47acef93-28e0-467b-8706-75b236f882bc · outbound

This paper cites Nock and F.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and F

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.713336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.854967Z digest=sha256:1a9620347190efca385e2419bf26b40ee6469e95a9a3d88934d2b3bd47935b5d

Observation 19057d04-3f49-4fe8-a5dd-af41e8ea68d9 · outbound

This paper cites Nock and F.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and F

Reference 39

Resolution
verified exact
doi, observed 2026-08-06T18:52:38.889685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.915265Z digest=sha256:2efb961db7aa406281852706eb8fa3926b410e450f8b3afa32da0cea4a38f390

Observation c59ab0c2-4ffc-48e9-ba54-f641f9820612 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.699922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.001364Z digest=sha256:d02924384bc46dbf631ada202c5c529b311ed64100b05185de38fc5b1b413a60

Observation d48ec375-fdcf-4214-8efa-b92bb651f903 · outbound

This paper cites Nemotron-4 340B Technical Report.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nemotron-4 340B Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.079112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.079112Z digest=sha256:faff3968f4e636e3010907979c5254707fb6cc442aa2db05d5998527ca678526

Observation 6d9a245f-f9fd-455f-8b32-49a6a3e425fb · outbound

This paper cites Ouyang, J.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Ouyang, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.687241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.152941Z digest=sha256:4b45884de71b0a9ec7a06287934bfba33d2d711568cb7c4a3eb910571b888a6d

Observation b96596a0-c6c6-45de-8988-40fcf3df7371 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.224508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.224508Z digest=sha256:e909b5c8769c50cbc0d2fc55c715e9e8a42fca859f10b7a864bd2ebb9a871ccc

Observation f56d7776-02e4-497b-9014-63d4212ff29a · outbound

This paper cites Rafailov, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Rafailov, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.673035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.280470Z digest=sha256:6df4f704e544225c189720ab8b5e8f7c8a9356e06d00449787ab8bfeda98bbd0

Observation c9a5e42b-2936-4a46-93f6-1f6c17925fc3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.659448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.389498Z digest=sha256:66045c80b786fbc28dcf275cf9c0c48664ffdb44bcaf1e06895b6836f163a9b1

Observation 4d5ac30a-a30f-43e0-915a-d8c703a5fb74 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.471398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.471398Z digest=sha256:75c20eba95cb522581b0a181521288d81e8bba948a20644f353381ad73b433c7

Observation cbb500e1-6af3-4c61-9eea-1934bd815d86 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.535852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.535852Z digest=sha256:4f8e9f742dfdc266ed23812e9046112480a3478554f4e8fede27521e353cbd0e

Observation 0d563e8f-888e-405b-992a-05b6de337c58 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.646073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.599101Z digest=sha256:3194c1d471bf6a5c81733a54c92a49f90821c7649e7c2e1c50dde0120a377ccb

Observation ac251eed-27f3-4bb2-8ba0-6b38c10425e9 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.633197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.665105Z digest=sha256:f1881a0f8c0d40a765acf87ac8a96222651f393fa0696e6e81efb91d46790da7

Observation 3ce4ba05-615a-4030-be25-2691ab56a28b · outbound

This paper cites Slocum, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Slocum, A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.619757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.723880Z digest=sha256:ed05f1f8994b295c653f904619ea0fb9b0bb8d927617feaf9c9e65f68a141da4

Observation 8812f057-e68f-4d3e-b688-476d9e86ca73 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.607057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.830259Z digest=sha256:c5f30eeb8dab68379f9ac07958ddfa077c64c34334cbd308b05ce2656dbfa959

Observation bd040ad2-ddcd-4ea5-8881-49e96b1d3e82 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.893399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.893399Z digest=sha256:822a24ce0ce93d8ce77bb477d92a173b36d9805ce3e5391cbbcca3e8ceb54025

Observation dac90df9-7ad8-44c4-9c78-95413555e016 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.594541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.954555Z digest=sha256:c6fdff51859d056ce4ab05e203e0a185ccd715c09eea260c0fedd193f2e33c3b

Observation 7f341fc4-d89b-4a78-bcab-887fe94274b5 · outbound

This paper cites Sypherd, R.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Sypherd, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.580769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.019369Z digest=sha256:641859043aca96ae1ec602c5d3c30c13e88a09cb07eec6672937ed8566b19d82

Observation 9daebe3b-71e2-4a3b-bbe4-99ffe76e8547 · outbound

This paper cites Tunstall, E.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Tunstall, E

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.566845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.147360Z digest=sha256:f3aa3beb30b1e30aabbd7c84ae5cf2d018c4412c3ccbbec1387400ab58ef3f7e

Observation 88ca329d-0891-469e-9bfd-9ba8e26cdf17 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.553032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.268692Z digest=sha256:a0d78570f5193a9f98d244a5e3235f1015646263a915bdbbdfdcad737fdcc034

Observation 50fd11cb-35ba-45bb-bacb-9fb4b9a50010 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.538286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.437209Z digest=sha256:bf6887efbbca2c0ba4bc4d85b61f2586442d94194f286a6a4cba4885b52c04c3

Observation bfc3f260-911f-4edc-ac4c-2c45d148e19b · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.524264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.563817Z digest=sha256:b398684d6b1f956d26ddc601d7a9926c9197b618bd15c8104ed1359d47bf370d

Observation adc61e49-10df-49f2-96e6-7f0df74fd2f1 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.510989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.682104Z digest=sha256:8323b7ff482408a1d73b8924a15b35df4a5055b4c0ea84f9d232468e2f81d0d1

Observation e22ddd89-11b9-4ec6-8db6-9870ea80c6b6 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.496283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.778837Z digest=sha256:6fa4cc6656c5249ce1c07c638f5d07ff42fa903302b5ef7b34d1709a4207d354

Observation 08a406d6-ea26-4385-ad56-1768411060d0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Qwen2.5-Omni Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.783043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.783043Z digest=sha256:1482420e9f6810eeba954eb93e296a73416956b60de47deecdfd74811b54fdcb

Observation 87803603-1cad-408c-b14e-b59054fc6f47 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.480895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.787672Z digest=sha256:958baab29f0080856ec4d0b545b8489da58b89f31d5c803b0923b8bd03c4a90a

Observation 0881ed7b-d103-4605-bd91-21ebd84de778 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.465569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.791473Z digest=sha256:e7f374885ddf4bb8a38cb16155052a177b0c82c79ffa198ede7417d3cb4a244e

Observation ec96c192-f2de-446e-a537-58949b0618cb · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.451379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.795690Z digest=sha256:96b8aed9a125e413f24f6b07d8dea73f24345133d6ac620238196a246446e3db

Observation 87ebc61f-baca-466e-855a-47b090309761 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.799899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.799899Z digest=sha256:73968d7580ec57c00af4d1d06ec78ca215d8e13ad37fd629b05d4cba03963d3f

Observation 57dd9b03-2a44-46d3-b077-a2ffcc84e3d3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.436613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.804815Z digest=sha256:f097ec5a024098df7926440adea659d4816bcab6eaf35f1eb69388f01de69af1

Observation 6b833df4-92a0-4977-88b7-29c850e08e7f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.808890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.808890Z digest=sha256:587ac57bcebf950ba49ff2113abd95761f737ae1c3fc90733e364294355fb57f

Observation 0e55d8d0-b384-45ba-917f-38e4578e5718 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.421914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.813440Z digest=sha256:7ee14f3e00fdc8621baac9c823ebbd6917456206ca7b9a008c7c749973a3cdba

Observation c8c75fc2-f3d9-4982-935b-29a743119ac3 · outbound

This paper cites @esa (Ref.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) @esa (Ref

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.817370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.817370Z digest=sha256:0dec348d64f396649ec1f355eaa1dcd3590cad27b54bcf0d0ba34be56c745596

Observation acb13461-2315-4188-8636-15f03c2cd226 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.822165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.822165Z digest=sha256:73d54b7123240aecdc5ea631dd610f6f55e5944b4d409a8eb639fcb11d138784

Observation df42e405-82d6-4b7c-92ba-e6aeddbf3765 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.826438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.826438Z digest=sha256:2ee76917ef6fbc60e97ce653d4b527916a1facaeaf151a90bbc25e9926940fcb

Pith citing papers

No inbound Pith citation observations are available.