Pith. sign in

Paper Citation Record · LEDGER

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2501.03271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03271 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:52.822369Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b8df58e-5994-41f7-8a91-f8e5c8f778b3 · outbound

This paper cites It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.423899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.660007Z digest=sha256:5727da43973111d27c7a3d3ebb653e278bdf2b2ba78e3ecb503605453a59a8f8

Observation fbe7a981-3107-4b70-90b2-cae9a7b39673 · outbound

This paper cites The factor γ determines how much the model should focus on aligning responses semantically.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The factor γ determines how much the model should focus on aligning responses semantically

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.413752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.663766Z digest=sha256:68b12a045bdb0daf45f739be520469c4ed5d27a2aca3f723f7a120038c6e6980

Observation faaf65be-ec15-492b-963a-a6dab4c232d7 · outbound

This paper cites semantic margin.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization semantic margin

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.384130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.674734Z digest=sha256:e8de67742e6b0d510b200f98a176f95b53ea4aba916fd6ea13fd119bdbc6ce54

Observation 39713119-a846-4b51-b765-999a12a5242d · outbound

This paper cites Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.320279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.696282Z digest=sha256:5a79d6281eae727dbbbf6901328e23f9419ce5de49e155a8c09f7eaca0c5abed

Observation 98083ef1-ae6d-43d4-b938-84e025f9787e · outbound

This paper cites Advances in Neural Informa- tion Processing Systems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Advances in Neural Informa- tion Processing Systems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.443777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.649154Z digest=sha256:19f7e21ac6d9a221e7bcbb3d8a567b05db0f2d2f8570362685a87e3ec5f044dd

Observation 1a6a632e-4802-44c0-86a7-9cb8366928f6 · outbound

This paper cites SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.652196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.652196Z digest=sha256:f038aeae869de0fdfc9007d6b5108fc69ccd5d8ac1b3c2ca4f490811a2fef56e

Observation d0d48f6e-9af0-4a58-ad8b-d144d1b7fe4d · outbound

This paper cites So the answer is,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization So the answer is,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.433940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.656680Z digest=sha256:4361e30c8119a5fa4ed2046aba2e1d5fc691c2c56f48b98e226ee58c83e14042

Observation a2251d4b-5fa0-41f6-96c6-f4bddf191a58 · outbound

This paper cites • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.404798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.667668Z digest=sha256:f192c7976b0deae48bc53afdeed015c295ca6ef4b5cd99a406bfb5af24e4272d

Observation e8dd7623-0101-4704-b1e8-3310ce513e25 · outbound

This paper cites This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.394905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.671293Z digest=sha256:3df214215839babfdba944ae05f11bf128d42fd7d713a15e35229af99e2e6938

Observation 9688a092-39bd-45ca-9740-a25b6a7d8856 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.373753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.677865Z digest=sha256:8603513c9ee00cbf8441b5969b563da8cbb15d4d50bb9f0e701ca415479bd860

Observation 85cbe184-9c23-47ea-bbed-ee6722f31e0f · outbound

This paper cites the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.363449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.681179Z digest=sha256:48dbb60f1f2ec423a5610c87413dd21d70b51876ac6840986f18f0e046fd2d88

Observation 5e404d99-e947-4833-8ecb-a24e47cb5d7c · outbound

This paper cites It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.351871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.685421Z digest=sha256:d6189470bb72ca15992f237119466844fe1d57cba4c86577b9a53faf9ca383ab

Observation 483f8a42-42ba-44e7-9546-f1f423ca072d · outbound

This paper cites Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.341110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.688857Z digest=sha256:3760ea6f80d77c4260e5490b5cae5f39c8ea75cfab2d08eec5e0a6b5d669e6d0

Observation 04900dec-2822-4706-a463-90d00936f45f · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.330517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.692931Z digest=sha256:a8daca46fe0c96d6a902a642fcff16ffe27065b56408347c41d391787d3ceee6

Observation c151bc98-d2e7-4869-bef6-cab0287aef28 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.310912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.699586Z digest=sha256:1a7565099f8b5ede39fd60b201e201070168e37ba07729e7dfc8317926c46465

Observation d223f0bb-a233-4dbc-b864-7f5968d2e625 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.299967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.703029Z digest=sha256:b095cd21cc32baa0ed7adcb7ab55c6c3790ec132867747d9ad608a7b09560bd9

Observation f8c00141-ce8b-4703-8ae9-fef6093ab31b · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.290903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.706207Z digest=sha256:e7f4f09a60b846255a9c7aad68cee67ea1da632084b9707b2b1f09211fa56b63

Observation 2eb72395-f5a7-4f3a-9dcc-e5abccee9398 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.280277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.709623Z digest=sha256:aa877ed2b3515fbc1c00b52291378d240e4725db6f3d73b642725023529667de

Observation f21edbfe-7652-4769-b057-5f057ada81b5 · outbound

This paper cites The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.268810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.712628Z digest=sha256:47aa89d6a405acc49ad5ff4b682b532b90e422456170eadce092f1b9407d992b

Observation 5aa28a01-606a-437e-a347-f81020c21f8c · outbound

This paper cites local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.257548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.715537Z digest=sha256:910c95faecc312c608fb4c11e4b07a53f67bd9a53a33a5b1389f3a5e5e0229a3

Observation c6e680dc-33de-4d25-bd4c-fe538c67ad57 · outbound

This paper cites • Computing the logarithm of the ratio between the positive and negative class probabilities.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Computing the logarithm of the ratio between the positive and negative class probabilities

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.237843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.722134Z digest=sha256:b8adbfebeda1183710fbee9c5ada81f81884f19ef4ebf4144e73b987d1be7a74

Observation 5b876d01-a45e-43d9-bdc3-6d3c931ec274 · outbound

This paper cites (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.228407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.725251Z digest=sha256:ee1fb09739690c556d3f10a7b2e95e515d362579150332d1830181ac816b942a

Observation 654f3029-6129-4f9a-b50d-13332c6bdf9a · outbound

This paper cites • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.216828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.728455Z digest=sha256:c14151146222fb6455d683862c872353837c6dbf635b9d88cdd63594798c63df

Observation a8b48bb8-0205-4ac5-8c9c-887316c40714 · outbound

This paper cites − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.207555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.732117Z digest=sha256:c53dd2c5c495ea089e7b18416318f3ea76d9a3030ccbbd70d7d45d34dd34f633

Observation 769a11e7-3947-472e-b436-a357e4b65815 · outbound

This paper cites where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.197893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.735984Z digest=sha256:3a0b7df4991f34c1cb168568ceca2309bfbd8c9af371a92b3d2e88c3c10b365c

Observation 6f955f84-b230-4149-ae58-a95c08e763b5 · outbound

This paper cites • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.187809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.739234Z digest=sha256:cdb61cd53c607edf15a3720074db961566afab74cfd81e63b8392c74c1ea19a5

Observation ccfdada3-6bd9-45ed-a461-af5c88418575 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.178490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.742997Z digest=sha256:aace13aa1707889f54b20000afb6076e0afb29f6a22b72c9ffed78ec70dbb269

Observation 58cdebd1-44ab-4239-8527-29214910d430 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.168497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.746872Z digest=sha256:952c358358acacf09ce6684c8ef177e0ca4a8708aaf12382d45e0b649172829e

Observation e4db9c6b-6be1-451f-a2f2-f5a1c6209108 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.159314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.750944Z digest=sha256:4efd2e7ac76ea0f0be140f5405eb4edbdfbeb6a8716b0dd55da39e27e3b3a385

Observation 5c4b5597-7c17-4547-af11-348b02da768e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.147808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.755229Z digest=sha256:4ae111d55e37d1aef97e1cf0ab37a0440f6572df6ead2b992476b3f46d9202ed

Observation 31a46170-d363-4e13-b13b-0d3e0638337e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.138212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.758966Z digest=sha256:2cb980c12905903be087ccc43e1c4a51b08d66cdc26d8fc401f5e2e817ca543f

Observation 128b0f3d-9da8-4dfd-8f30-9863717d3453 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.129088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.762517Z digest=sha256:0005176f0b535b28e0995e32d47377e1f25a113bc7fa93d40eb4ff69498891f0

Observation 5eea5219-43ca-4e1d-a53f-2ded0c5a691c · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.119709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.765952Z digest=sha256:e04e658d4ff354430a91015a1cab26f5ef5ea903037a5afaedaed784e0b10441

Observation eb5ad9ae-ef82-4688-8a1f-2496f3fde62a · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.110711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.769188Z digest=sha256:dcda10de1b3d0ff6f643cdd62bd66adba9fcc561e0f8798fc681d96ac2bb42f5

Observation 860061fe-95bb-4ff3-bbca-a4e2824bad87 · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.100371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.772893Z digest=sha256:3dd927364e5dfd2412509f0717ce5e655e4294ea2637f194e96151adeade48a7

Observation 7ce2653e-f1c6-4cee-95a5-3148b580bac6 · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.090521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.777240Z digest=sha256:fbb528ddddc3991de6be4092e3ee0d2ea836c20db9b7fef9c003f73a36687f96

Observation 46f1883d-dea7-4ef1-9e4f-1611435d13ac · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.080425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.780784Z digest=sha256:123471209700d44e81a13de07b1acf494901e75725f15693eb2f68116f77d1d9

Observation 4202698c-b24e-4ee2-8a50-34fcfa2c643e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.784682Z digest=sha256:7dceee00727b8769534e0577cbe762220d3bf5d21e6f81d9969196d8b577144c

Observation 90bbd48b-8d29-464b-b884-a3e9dc6ffc07 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.788403Z digest=sha256:2ddd025965543a94d6f91730b7ac6cb6e6e4f5b8ef1d13f0a8e13cf87b707611

Observation c80db629-ddcc-42c7-b1d9-02ef01cf52f5 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.048614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.791496Z digest=sha256:d6aff11076d09918bf5850bc6322c1faf55e34ddb927ca400670992162c394a7

Observation 84a73a3e-fa9d-4902-b819-a7bed2d439ba · outbound

This paper cites RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.039153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.794313Z digest=sha256:7f6184a9c2919ee26c715380e80a7b65ee63629ff47401515246bc1fd31dcb34

Observation d58bd280-5e59-4bc2-bb9c-7b545da80ca6 · outbound

This paper cites Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.029184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.797742Z digest=sha256:db2c867c0aa8a8b60beb5120eab95e21943b6b42aa7d3f714ab637f7f294512b

Observation 199e6983-76aa-4b0f-9b1c-0edcdc30b325 · outbound

This paper cites • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.019218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.801100Z digest=sha256:47a168e22ed836704a61edeb4efb5f97d6b5c05526f127a00c84d00d596ea9ba

Observation 96c4d69c-0fb8-4d61-945e-57d561595d32 · outbound

This paper cites Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.003717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.804028Z digest=sha256:6f62e792545d2b007ec3431e0b47149b4b221d30c67ce2afad170222920566a5

Observation 2b6085b4-e054-4701-ba60-6a1f7ead3b5b · outbound

This paper cites Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.991567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.807196Z digest=sha256:5603aee3fc0365c5e622cec9ed90a344ac7c3051443585e417d27bead62cec1c

Observation cc4920d5-3a07-414c-bd63-6218ca5b34fa · outbound

This paper cites distance.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization distance

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.980602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.810821Z digest=sha256:9c3f12d6129d324a37f8290b233a175fc398456cade5d2e4e196a4b2b3ee49ed

Observation b85bd9b0-1645-44b2-8d4d-ab4e3fa171b8 · outbound

This paper cites HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.970574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.814975Z digest=sha256:399c98f46187aabeb36b4b3e9d4b69149692347910c02d6434918a12fb453a40

Observation 06cd8f15-93eb-4c3d-9565-45e359970603 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:52.958609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.818646Z digest=sha256:2b4beb8e5634d64264b5bef87c7ab0df33031607ccda06dc9a18cd0bfd79da10

Observation dff8693b-f982-4353-bb16-a2fb432f6df7 · outbound

This paper cites Correlation Flow,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Correlation Flow,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.948103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.822369Z digest=sha256:e8065c873a7e12091cd7b33807a8d2045727faa2eb909e38690529dfb41e12aa

Observation 627141cc-c26f-4354-8f0b-063b89ed57f2 · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Towards A Rigorous Science of Interpretable Machine Learning

Reference 465

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.616572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.616572Z digest=sha256:56a18b69260b3bbffef37f251bc2f48aed2b8e721f93596d4eac446b997b156b

Observation 4d2e9a84-5d0d-4993-a114-52a94c6f4b7c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Representation Learning with Contrastive Predictive Coding

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.636195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.636195Z digest=sha256:3c4972d48313edcd04adfcf3eaf398c5ea8a57c5ab50b4cf34399a653e5737a3

Observation e8ab9796-9968-4d86-b365-93bab8e9ef69 · outbound

This paper cites stop execution if X is true.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization stop execution if X is true

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.246861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.718765Z digest=sha256:0f99926ae1eb16e1770e0ee4b2985fc2f373029350728699544d877497ba1a20

Observation 735ef1a2-1143-4332-8459-25218799a086 · outbound

This paper cites In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.463425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.620948Z digest=sha256:00607bf11d1306e767cc98ff500ecf8f53677b336f50fa1c0692c2beaa01768c

Observation 161117fc-6c47-4808-bf29-27efd01ab84c · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.624338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.624338Z digest=sha256:42c43c86452dbaa296bd90221b53fd3c40289b9ffacc85b23e8df78598e8a762

Observation 8efc0ae7-d759-4fe7-b41a-a83990e3f85f · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.646022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.646022Z digest=sha256:3a4e9f98e71ebecffc34e50787bb0964c98ee3451171d2e919b851f74905c937

Observation c7c61bca-92d7-45f7-a2de-b5a16724f374 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.643057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.643057Z digest=sha256:e1226136fdf554e2e6f1ffbb2ac70e1d27e90072a8fe9c72f271074d2504240a

Observation 36d39408-377b-4298-ab49-eb471f7f3c6e · outbound

This paper cites In International Conference on Learning Representations (ICLR).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In International Conference on Learning Representations (ICLR)

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.453574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:19:52.632670Z digest=sha256:ca83ff8708422b187db3f0db6072fbd6982777aa9a223c44e1c608a05ff757c8

Observation c685b90a-17d4-48ae-a7c7-1b2dda82f10a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.610934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.610934Z digest=sha256:59ef8d62f208c014df59dd1a297cd1623bbde1925316f4a23f808d99fe3122a2

Observation f71bc17a-6336-4ed3-bc88-ea28b62fdd03 · outbound

This paper cites Training language models to follow instructions with human feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.639456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.639456Z digest=sha256:33b29a0c145116a4234e2e46871a3240d18069922f3a1a432ab5d38f10a41743

Observation 3e278364-c54a-4a64-b026-deea65b7dd89 · outbound

This paper cites Let's Verify Step by Step.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.628333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.628333Z digest=sha256:6e2808ef12dff906528f17784596ae2bed1e953142950c89ff2566def0c5a82f

Observation 32c14187-a3d3-4346-9292-90415eed753b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.606670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.606670Z digest=sha256:fb5a670930b71511118b8aa92ca8484765d4b3ac078ed456009bddeb31cfd079

Pith citing papers

No inbound Pith citation observations are available.