Pith. sign in

Paper Citation Record · LEDGER

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2502.04567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04567 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:27:11.573917Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:58.461665Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:36:40.511807Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09ff3f3f-553d-4098-9317-addc4c038319 · outbound

This paper cites G., Guo, Z.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator G., Guo, Z

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.391147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.391147Z digest=sha256:3a33b8810e398045134d10220b0a34914d6a095ef5993af3b0dd75d364a3ae2f

Observation b7f98236-acca-44d5-92ad-5d844804e912 · outbound

This paper cites E., and Nocedal, J.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator E., and Nocedal, J

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.396046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.396046Z digest=sha256:39f3d3c2d1a08dba3c1e6bbbac1bebe09644607b4c1d905ae732b4c697b3bfb5

Observation 5467d72f-8252-4146-b306-2a0d4123fe8a · outbound

This paper cites an unresolved cited work.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.402075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.402075Z digest=sha256:bb2faf90087bc188bf54b5ab44f73cb1054449858dc7786f63ac76ad8e4bdeda

Observation b4898ca2-5e47-4c0f-a946-6fc3ac606d25 · outbound

This paper cites Noise contrastive alignment of language models with explicit rewards.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Noise contrastive alignment of language models with explicit rewards

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:12.042617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.406347Z digest=sha256:1e575aee9da47fd40f5f4bfc25499376ae96bca0f51dc70887ba958f151fdcdb

Observation 7346bb9c-6aa5-488e-b262-cc6f8f0b2e1b · outbound

This paper cites Towards Improved Preference Optimization Pipeline: from Data Generation to Budget-Controlled Regularization.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Towards Improved Preference Optimization Pipeline: from Data Generation to Budget-Controlled Regularization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:27:11.871651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.411455Z digest=sha256:308eb7533b10278fee97aac09fb965ccbe6c67aefb6bf560c17a9520220accb3

Observation 75cdeb0f-4ad1-4c10-984a-685cf5c4f2f1 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:12.028948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.415943Z digest=sha256:64e12ea35dfda566bdddb1ee60e140b99dcb6bcb2e86895b327cc1fb567a3bc3

Observation c0d6ab18-1fa7-465d-b33f-6b2767e23a04 · outbound

This paper cites Anchored preference optimization and contrastive revisions: Addressing underspecification in alignment.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Anchored preference optimization and contrastive revisions: Addressing underspecification in alignment

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:12.014457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.420602Z digest=sha256:66954270d62a98482de09c965a1db5ad009abfb7d730ba32cb1afa6ad974699a

Observation 39445c87-3d7c-40dc-b8e9-7e36c312a16c · outbound

This paper cites The Llama 3 Herd of Models.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.424913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.424913Z digest=sha256:07b98c14c20745ba33e6c3a4594432b1fa1d0fad55f3fa83716c1f580c01afc4

Observation dc34f322-dd1d-470a-bd37-86f452694010 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.429106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.429106Z digest=sha256:c86465ed0c4e03d09e7d087dc70983a611cff02ba7cc175f0e4b13239cecc445

Observation 88f3c83e-1990-48cd-94d2-8d307d1c65c6 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator KTO: Model Alignment as Prospect Theoretic Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.433240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.433240Z digest=sha256:113553c788849b1970b832a792fe9fa93cbf36fae0e6e9408986469cff87be2e

Observation 2dcb33f1-8483-413f-b250-8e4b0ab79b67 · outbound

This paper cites Aligning language models with preferences through f-divergence minimization.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Aligning language models with preferences through f-divergence minimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:11.998153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.437747Z digest=sha256:89d953124b53b36b002e443d6ba882f831ed5925cc294b108ddcf7428805eeb9

Observation 35a8d537-a80e-4db4-a8d6-ec6e39ef4f5e · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.442396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.442396Z digest=sha256:b64a75cae32bdae12c2e9ac866934330d9a6bfa5a79c9abac104aeeeb9c06cb5

Observation cc708a27-0e70-4be8-a02d-77803f193388 · outbound

This paper cites and Hyv \"a rinen, A.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator and Hyv \"a rinen, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.446933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.446933Z digest=sha256:634d43215db4269b0ede4c9d78842cf13743b43f864a0830fda6dc54a92cf327

Observation 54e29e08-df1f-4a04-8eb8-4cd62e0c91e4 · outbound

This paper cites an unresolved cited work.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.450635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.450635Z digest=sha256:b0b1bced0e31e7bcb339c4faa344ecebf56d9ef64e335d4bacfcdd2108808bb1

Observation 50f27505-97de-4b92-8df5-5e535e5f01c2 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Towards Efficient Exact Optimization of Language Model Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.455229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.455229Z digest=sha256:17ba2c43cea2b9662f517c12768a2bc28ef0a3a3fca7fd63ba6501a197536917

Observation c4d54cf7-f402-4d13-b4fd-a5817cdb209c · outbound

This paper cites Binary Classifier Optimization for Large Language Model Alignment.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Binary Classifier Optimization for Large Language Model Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.460376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.460376Z digest=sha256:c54d1dd743d05c914e78406d2a59be91e7b8d916a8de1ec2b63da1c2a11eca84

Observation fbe46b0e-7036-44ba-9b12-059c391530f2 · outbound

This paper cites H., Gonzalez, J.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator H., Gonzalez, J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.464975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.464975Z digest=sha256:b496a413a46cdbe69072833d4f285cdd80c3d1a778ffa8e9185e8da655a59500

Observation f3b3714b-1694-452f-b7c2-092cb8ee9d0a · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.469519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.469519Z digest=sha256:21349600a712cf06066a12c79e3c8d0931c207492e54628355a49e7cf66b51cf

Observation 05a7362e-cc12-45d8-a1cf-48eb939c46a1 · outbound

This paper cites E., and Stoica, I.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator E., and Stoica, I

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:11.959427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.474148Z digest=sha256:fdd14eea4263ef9b0245c0a97570af3d59abc6b7b2be2a279f32995b2c8e6fc4

Observation df2d7f18-845b-4a78-b1e6-f792ebd39473 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.479611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.479611Z digest=sha256:0c6796916ba65bc39d1a188822acb7cec742cb4b8a5d944bb1818f15dfd307c3

Observation dad8a705-5d0e-4ce1-bad6-2c9cfea9a834 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.484390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.484390Z digest=sha256:d69510ce2c13537597c221f20a79b02a2fbeaef2e59d290a468d6b4b3baab58c

Observation 88278b66-2fb6-4044-b725-e202895d8152 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.489037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.489037Z digest=sha256:ebdb0d61c1b2a982025cb934a01027e32e8582054df1d66eb8ec1d49bb8a2b91

Observation 0f9d9fca-3083-4181-b1cc-b7ffc7a440d3 · outbound

This paper cites Elements of Sequential Monte Carlo.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Elements of Sequential Monte Carlo

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.494453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.494453Z digest=sha256:d25c52b7a3910de4462f01d511779a55b38b3d395d5a50121b2a4889a1811d30

Observation d44bc280-46d5-4d8a-a633-4e725dfccac6 · outbound

This paper cites On the connection between noise-contrastive estimation and contrastive divergence.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator On the connection between noise-contrastive estimation and contrastive divergence

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:11.946431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.499212Z digest=sha256:3fd26549206fc864d675f237741bc4c6eba8028e123b2d6e865893aaec5ca524

Observation cc703708-848f-4776-9c6b-79de8ddd461e · outbound

This paper cites Training language models to follow instructions with human feedback.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Training language models to follow instructions with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.503738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.503738Z digest=sha256:090df4d098d48bbc3850db60ad921d2af36db93a2e96638b9b8e297281408fe4

Observation 4ad16160-6c38-41bd-bbcb-c7b5c618f451 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.508003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.508003Z digest=sha256:604b290654dc983320c941d737b3624a565a30e1a1c4406d08e276848b8247f3

Observation 667a8f3b-64e8-4391-a4af-d0f26ba1f572 · outbound

This paper cites and Schaal, S.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator and Schaal, S

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.513501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.513501Z digest=sha256:df41e5fe15f3eb92afa0677304d5428e022a66a8460751778967a7d0fac5e377

Observation 14537d1c-3b73-45bc-a6d9-917caa9cc877 · outbound

This paper cites Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.518340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.518340Z digest=sha256:c99bfd19b94a4e738e4e2aedfce39e976375bcf3b484495b67477ef2450c9d9a

Observation 7d446647-df40-4f84-ba8e-3d50dc7770ad · outbound

This paper cites D., Ermon, S., and Finn, C.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator D., Ermon, S., and Finn, C

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.523367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.523367Z digest=sha256:0e15e4f1dc183761ce63cbd4f7243d0156b3c1e6ec2568e136275162d5cb382e

Observation 95964de6-65b9-4c72-a694-e5fa675a7a98 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.528056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.528056Z digest=sha256:6eb97902e6af3aa4ff4f53ff37201ab2deaf7f7681d7fe5a471b326df68bfd03

Observation 5e346559-95c2-4063-973b-b7676843897e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Proximal Policy Optimization Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.533210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.533210Z digest=sha256:40b0631271e0fa89777092cd835133a961cae55cee6ae310a2626bc83ce330c8

Observation fc894399-ce8f-495a-af70-8e1f93fa99de · outbound

This paper cites an unresolved cited work.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:27:11.909436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.538375Z digest=sha256:b2bbca6faef33f22375c669e91c763bdf33ca116b3b9c8620cd987f88243f8c1

Observation b788fd2a-d719-4b25-9826-3906610d7ff4 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Understanding the performance gap between online and offline alignment algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.543782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.543782Z digest=sha256:32abba29a151d477727dac41841cf509c43a66ae1175bff69354bb79f80dd31f

Observation 94169a91-f4f2-4441-9332-0d867585e7bf · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Zephyr: Direct Distillation of LM Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.549336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.549336Z digest=sha256:dcce98bdd6370a8a47dda51a4839f47f3ac5fab49b7662246fd1b9c5a82238e0

Observation 30a80611-c1ba-4d9f-8571-f9768e024683 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Self-Play Preference Optimization for Language Model Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.555427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.555427Z digest=sha256:66b311e836f9474acae205565b26da7e35709bff960f7f5e9a00c7a16028fe99

Observation 23606bc7-2505-4ebd-8220-6702321e2409 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.560059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.560059Z digest=sha256:21a105ed2c3c2ac445ead73f5e6854c8a4989e31cb437acc776ccecae5a3cca2

Observation 88fccbc1-22cf-488e-b8fb-5bc9d8170d15 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:27:11.895232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T22:27:11.564541Z digest=sha256:e316d65030059d8bb7430c65d105c43262461279d2c49a3c3ebed07ca668e66a

Observation e9d8cecb-6904-4a76-be1a-5d3f8f3a8008 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator Fine-Tuning Language Models from Human Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.569074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.569074Z digest=sha256:547d8a8089312051521502fd36a6913b180821f0308ba13ba342d62043348712

Observation 5528dc08-a4d6-48f8-9b59-19331eba3f92 · outbound

This paper cites write newline.

Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T22:27:11.573917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:27:11.573917Z digest=sha256:6e8249d0681613a5feb3b5a7f016fc45f5d32a3f9d7ce532633dcb882e032a39

Pith citing papers

Observation 8216cd18-29cc-4630-88ef-433fb7559ef1 · inbound

York's Cavity Formalism and Quantum Modified Thermodynamics of (2+1)D Black Holes cites this paper.

York's Cavity Formalism and Quantum Modified Thermodynamics of (2+1)D Black Holes Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:58.461665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:20:58.461665Z digest=sha256:67cbea06152cc36383a8a4320b5a6024e22b4684d80ccf5841047217734a722f

Observation 268757a8-56ca-4ae9-bb86-dfe3e9422386 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.524459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:111ad354b2e5c5073137fd47873189d408ec768ae2fb1aeafa490178cf1d4106

Observation a4553154-f928-474f-9534-aa324d786347 · inbound

HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation cites this paper.

HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:40.513620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-25T05:30:37.320976Z digest=sha256:1f1c441272f50e76e952558f43799d45e5a0a60461ca0261aa6eede10226a609