Pith. sign in

Paper Citation Record · LEDGER

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2501.03271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03271 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:52.822369Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b8df58e-5994-41f7-8a91-f8e5c8f778b3 · outbound

This paper cites It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.423899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.660007Z digest=sha256:767167fb9872b4d27ed2194870a38308f176210d6c2402df4845976bc6921fe5

Observation fbe7a981-3107-4b70-90b2-cae9a7b39673 · outbound

This paper cites The factor γ determines how much the model should focus on aligning responses semantically.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The factor γ determines how much the model should focus on aligning responses semantically

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.413752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.663766Z digest=sha256:7f12af9d188161ebfee04bfa192a54ac17c49a5953ed3725f4abcdaeec5a4c33

Observation faaf65be-ec15-492b-963a-a6dab4c232d7 · outbound

This paper cites semantic margin.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization semantic margin

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.384130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.674734Z digest=sha256:8fa6df33e9b041898780711d88b4047f3ca460dbe05ec5e9521cb87973216425

Observation 39713119-a846-4b51-b765-999a12a5242d · outbound

This paper cites Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.320279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.696282Z digest=sha256:ec84c33687134755c1443c92239cb5a1eb114802211cc87bd40014a614a00a2e

Observation 98083ef1-ae6d-43d4-b938-84e025f9787e · outbound

This paper cites Advances in Neural Informa- tion Processing Systems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Advances in Neural Informa- tion Processing Systems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.443777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.649154Z digest=sha256:f3e52eda4fc120e4deb49a4ffb91de59f10f15caec3fc47f2eefbd6956b13dfa

Observation 1a6a632e-4802-44c0-86a7-9cb8366928f6 · outbound

This paper cites SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.652196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.652196Z digest=sha256:373acbcf9ace6426149bbd45504e3ff51d6d472527056c9319d9b718172bd949

Observation d0d48f6e-9af0-4a58-ad8b-d144d1b7fe4d · outbound

This paper cites So the answer is,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization So the answer is,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.433940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.656680Z digest=sha256:34336168439824913e0932bb0430bf96b08a7d308ba826d5e12fac58f4968611

Observation a2251d4b-5fa0-41f6-96c6-f4bddf191a58 · outbound

This paper cites • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.404798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.667668Z digest=sha256:c4490707b82f64156e17221ff04b04a91ceef1d5530d81d8d67734aa5fd358d6

Observation e8dd7623-0101-4704-b1e8-3310ce513e25 · outbound

This paper cites This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.394905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.671293Z digest=sha256:e968c10548d0d6eb81a0d51975b2113da6011fbfce9e3ec08f7f5336249eae24

Observation 9688a092-39bd-45ca-9740-a25b6a7d8856 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.373753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.677865Z digest=sha256:3fa71a88beab78dce029b2f99fd07f84f49d5850605ebe2e1548a722a3b11d54

Observation 85cbe184-9c23-47ea-bbed-ee6722f31e0f · outbound

This paper cites the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.363449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.681179Z digest=sha256:388aebbd309525ea2310ff4f19d335ece5e20ec38e5915feb2ef673c66bceb07

Observation 5e404d99-e947-4833-8ecb-a24e47cb5d7c · outbound

This paper cites It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.351871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.685421Z digest=sha256:3168f8dff27e52b1ded7aec350f0c0d9ff62abd686e5b7855b64ce1a4c4a0116

Observation 483f8a42-42ba-44e7-9546-f1f423ca072d · outbound

This paper cites Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.341110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.688857Z digest=sha256:3a45181e9cc6f63349ed80b110059327a4670774ce6249931f72dde8a45b27af

Observation 04900dec-2822-4706-a463-90d00936f45f · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.330517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.692931Z digest=sha256:d819b33bdf7e2ae280b864acaecec9240aa7b3efc00317d04ffef746a7cc0138

Observation c151bc98-d2e7-4869-bef6-cab0287aef28 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.310912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.699586Z digest=sha256:23e26fc2387c82cd74a206d4d1e7c795c5c8d165ea97cd995ce87ca1bb389d3d

Observation d223f0bb-a233-4dbc-b864-7f5968d2e625 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.299967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.703029Z digest=sha256:02ace79849f7724e88e8c9c786fd1e1376ae88b097fb1895338d7c8c147b8f89

Observation f8c00141-ce8b-4703-8ae9-fef6093ab31b · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.290903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.706207Z digest=sha256:8d2f37659c043bf269c80f688cccbd31699ae7160c9f03878dfb1de3107506a1

Observation 2eb72395-f5a7-4f3a-9dcc-e5abccee9398 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.280277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.709623Z digest=sha256:5be107d55b6831a637feda3545138df0748d6e186eb0bef1fd6f7ddca3467254

Observation f21edbfe-7652-4769-b057-5f057ada81b5 · outbound

This paper cites The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.268810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.712628Z digest=sha256:64ae6e8e972acd93c707ea5fe002d671af213d4bd1059e3423c5280ac5ee958d

Observation 5aa28a01-606a-437e-a347-f81020c21f8c · outbound

This paper cites local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.257548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.715537Z digest=sha256:d6f2efb1e6d6c387cfd65941d3f76220d6a883106fcfc176c1ef75ec506a3142

Observation c6e680dc-33de-4d25-bd4c-fe538c67ad57 · outbound

This paper cites • Computing the logarithm of the ratio between the positive and negative class probabilities.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Computing the logarithm of the ratio between the positive and negative class probabilities

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.237843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.722134Z digest=sha256:c95a354f6ca698c1f7356efd396a3c77b5bfc35c3ce59434970e92b054a04f4f

Observation 5b876d01-a45e-43d9-bdc3-6d3c931ec274 · outbound

This paper cites (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.228407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.725251Z digest=sha256:7e82e61a3c95b8b5b4cb81c20f0b977479723f620b1aee4d622ffcd6c6c8bf27

Observation 654f3029-6129-4f9a-b50d-13332c6bdf9a · outbound

This paper cites • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.216828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.728455Z digest=sha256:4554c822f0648356df33f003544b3c4ab791ab0d831fd630816d35988db7972c

Observation a8b48bb8-0205-4ac5-8c9c-887316c40714 · outbound

This paper cites − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.207555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.732117Z digest=sha256:2f9d52307dca3129409ffeefb3937b41b74229742f18559ecc538293a28ae165

Observation 769a11e7-3947-472e-b436-a357e4b65815 · outbound

This paper cites where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.197893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.735984Z digest=sha256:7e178ad31dcb64e6fae82a4ab664d6eaac12b8ddf04b4d00a22f5811921aca9a

Observation 6f955f84-b230-4149-ae58-a95c08e763b5 · outbound

This paper cites • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.187809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.739234Z digest=sha256:3d061a2f76240b4ea5ba5e698a3eda87fabaec6d39a67b82eba100e7d654de6c

Observation ccfdada3-6bd9-45ed-a461-af5c88418575 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.178490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.742997Z digest=sha256:60fef89e0e7d3f05d018ba0fb76be8dea03cfcd01ee9f6e98aaa5aa37c96bbf8

Observation 58cdebd1-44ab-4239-8527-29214910d430 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.168497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.746872Z digest=sha256:da039bf882ef1c1c910e3a97780cc64e6e278da346afecd47cebfafc2afe3b02

Observation e4db9c6b-6be1-451f-a2f2-f5a1c6209108 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.159314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.750944Z digest=sha256:bc0ba6e0c7c20391a3b5ed3d6349faa638d7666f2071842b82fca4b9101e88e6

Observation 5c4b5597-7c17-4547-af11-348b02da768e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.147808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.755229Z digest=sha256:aac2c87e5d14b1cf32ea14a98118cf4bc51e7fbf278586f5a44c4e0a65757188

Observation 31a46170-d363-4e13-b13b-0d3e0638337e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.138212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.758966Z digest=sha256:bb99b173e7b2efe3c2d547cbfdb140b99a33572bc47e4c4fc509fcd33433dcf3

Observation 128b0f3d-9da8-4dfd-8f30-9863717d3453 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.129088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.762517Z digest=sha256:a3eff9cee285758c6e90fa5800bce198967449b0875cc5869201ce79da0a2e8f

Observation 5eea5219-43ca-4e1d-a53f-2ded0c5a691c · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.119709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.765952Z digest=sha256:23b45ad3be128d726c9ed284021d9cbeedb85e6f65540165aa576c0efd76d51f

Observation eb5ad9ae-ef82-4688-8a1f-2496f3fde62a · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.110711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.769188Z digest=sha256:534be96da31e2b7877cdad9c7c72df90a36138becf884e757425421ef46ca677

Observation 860061fe-95bb-4ff3-bbca-a4e2824bad87 · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.100371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.772893Z digest=sha256:9e15873a3e8c4ddc5884d1b76e0c9c6e3d6459effc5a3e50fbffdf00fd8561e1

Observation 7ce2653e-f1c6-4cee-95a5-3148b580bac6 · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.090521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.777240Z digest=sha256:260363ee8aa05b481b40feb494a83bc20b1b405c4d252e1e63939ae0bb0ac759

Observation 46f1883d-dea7-4ef1-9e4f-1611435d13ac · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.080425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.780784Z digest=sha256:b9a76309f8b0cd1e4fa1bfd85fbafe1e61548b03fa0a045ce981d16d58383be9

Observation 4202698c-b24e-4ee2-8a50-34fcfa2c643e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.784682Z digest=sha256:8e7f8480ca2f96ff07f964b227e440d2a191bc160ed1929bb58705515c9c13e7

Observation 90bbd48b-8d29-464b-b884-a3e9dc6ffc07 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.788403Z digest=sha256:8048ff5a56d7e8421087d735544248288100d51eff9913609998ed32176664a4

Observation c80db629-ddcc-42c7-b1d9-02ef01cf52f5 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.048614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.791496Z digest=sha256:d90e7f6cbd602e4fe5444f4ff054e836f513df8d93aacda4b237f8f9ebaec6e7

Observation 84a73a3e-fa9d-4902-b819-a7bed2d439ba · outbound

This paper cites RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.039153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.794313Z digest=sha256:0b10b0f8019e536616dda07bfd272fb2f9c30caee68a0fea1738b3d81335d294

Observation d58bd280-5e59-4bc2-bb9c-7b545da80ca6 · outbound

This paper cites Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.029184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.797742Z digest=sha256:fee265543a5c23fffd6d5fb83ba67d4479adce3b0f9b7560d0dcfc47147a4769

Observation 199e6983-76aa-4b0f-9b1c-0edcdc30b325 · outbound

This paper cites • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.019218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.801100Z digest=sha256:95af0d6cd4740e322a1d80ef4a47ed9d9f48c8386b80cc77ee570f29a587159b

Observation 96c4d69c-0fb8-4d61-945e-57d561595d32 · outbound

This paper cites Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.003717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.804028Z digest=sha256:2752c32ece4e84a7817ca6dc1fce6844758725cbb9addc4a187ce745e1d535a1

Observation 2b6085b4-e054-4701-ba60-6a1f7ead3b5b · outbound

This paper cites Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.991567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.807196Z digest=sha256:17d8524e521e9f4a8448f86f565a727fd2d3049b15663de473c537c71d01b964

Observation cc4920d5-3a07-414c-bd63-6218ca5b34fa · outbound

This paper cites distance.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization distance

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.980602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.810821Z digest=sha256:05ee8879ee5021a4f0a5fc5909cabab698710276afdb6320bb329fd1a55a6a15

Observation b85bd9b0-1645-44b2-8d4d-ab4e3fa171b8 · outbound

This paper cites HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.970574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.814975Z digest=sha256:57731544f48c534ab2421edb171cbb677e1ffb96c0ffdcbfd3ffd287424a0003

Observation 06cd8f15-93eb-4c3d-9565-45e359970603 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:52.958609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.818646Z digest=sha256:8f32ad46d7955c4b1a31d8251c2fb1727198944572695bf9a8da422762ab550f

Observation dff8693b-f982-4353-bb16-a2fb432f6df7 · outbound

This paper cites Correlation Flow,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Correlation Flow,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.948103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.822369Z digest=sha256:0eb8c9bb433a4759de2e5b85e2c7c050f74e44e79d4bb829ab1fa965de8f3f91

Observation 627141cc-c26f-4354-8f0b-063b89ed57f2 · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Towards A Rigorous Science of Interpretable Machine Learning

Reference 465

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.616572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.616572Z digest=sha256:8d1b49d9025e61112ad60481cdc9a7f3e9afd41895e77809fe89e4d451ed426e

Observation 4d2e9a84-5d0d-4993-a114-52a94c6f4b7c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Representation Learning with Contrastive Predictive Coding

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.636195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.636195Z digest=sha256:5fa208ba981f821fdf3e4492ae75ca6f56c8b251f549a8e8af5839a33f8d1710

Observation e8ab9796-9968-4d86-b365-93bab8e9ef69 · outbound

This paper cites stop execution if X is true.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization stop execution if X is true

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.246861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.718765Z digest=sha256:d102526bb7499e2338f44502eb4ef002e1a83cb33f7d34219c73eace89f27c3c

Observation 735ef1a2-1143-4332-8459-25218799a086 · outbound

This paper cites In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.463425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.620948Z digest=sha256:9f19b2d4c722579c7a09636d1a49f14a5170f866087644d8028e5bd2b4e8005a

Observation 161117fc-6c47-4808-bf29-27efd01ab84c · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.624338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.624338Z digest=sha256:ea2881900ab5c23c7a4c206f23d6c0233b2bc89aace15f3e16fbe8f41b83db45

Observation 8efc0ae7-d759-4fe7-b41a-a83990e3f85f · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.646022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.646022Z digest=sha256:b920d24632b41e9cf59d7abf46bd6df213372e1f038000530aa4003b84cde1e2

Observation c7c61bca-92d7-45f7-a2de-b5a16724f374 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.643057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.643057Z digest=sha256:756db565894ee3aadd10415a4806c57221c64bb0a8d88855a970d0fb21038c5e

Observation 36d39408-377b-4298-ab49-eb471f7f3c6e · outbound

This paper cites In International Conference on Learning Representations (ICLR).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In International Conference on Learning Representations (ICLR)

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.453574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.632670Z digest=sha256:f552a500ef82086b127ee84be0617e3cab502e84ffe22442ff68f69e519b06a5

Observation c685b90a-17d4-48ae-a7c7-1b2dda82f10a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.610934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.610934Z digest=sha256:7c8196f3840f88167d4e542f4067b73405e62540d2019afaee4e60969ad8f269

Observation f71bc17a-6336-4ed3-bc88-ea28b62fdd03 · outbound

This paper cites Training language models to follow instructions with human feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.639456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.639456Z digest=sha256:7b09f13d5ee612ce71037f1d8117d221cd8226c4e116f9ad3326410b3cad5f21

Observation 3e278364-c54a-4a64-b026-deea65b7dd89 · outbound

This paper cites Let's Verify Step by Step.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.628333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.628333Z digest=sha256:d4e4c5a2ce0f16326264f55001b06648972fcafae56ae7541ccf135b5a8c2964

Observation 32c14187-a3d3-4346-9292-90415eed753b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.606670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.606670Z digest=sha256:5f37051d6715f0c1d92e15513c4ef9ebda740178550a6c3690e14543106939bd

Pith citing papers

No inbound Pith citation observations are available.