Pith. sign in

Paper Citation Record · LEDGER

Explicit Preference Optimization: No Need for an Implicit Reward Model

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.07492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07492 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:18.045861Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:49.149334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:01.276341Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 814c293d-8351-4c84-a266-071dd259de96 · outbound

This paper cites GPT-4 Technical Report.

Explicit Preference Optimization: No Need for an Implicit Reward Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.805167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.805167Z digest=sha256:325678505a9207db2a946c50543423425748650bbb52501b9eed41d06c4991ff

Observation 7733375b-fae1-45c0-94d4-4195aa3d2497 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Explicit Preference Optimization: No Need for an Implicit Reward Model Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.810011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.810011Z digest=sha256:9dd00ec39f39b6d9d18a1375ec6390e25fc6d0a45e6bc5e1fbe4f8b4f319c003

Observation 28a73e90-003e-4e7b-b6da-7f358d3c3f51 · outbound

This paper cites Llama 3 model card.

Explicit Preference Optimization: No Need for an Implicit Reward Model Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.814798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.814798Z digest=sha256:7d4979b2f2783839e5e0a4851c45338861d98a26ae22416328f075736923e479

Observation 587bad29-1ff8-4a50-9e75-684abf84c47f · outbound

This paper cites Direct Preference Optimization with an Offset.

Explicit Preference Optimization: No Need for an Implicit Reward Model Direct Preference Optimization with an Offset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.819727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.819727Z digest=sha256:fa4050c038c016730e3eaed9f66654d025c4a883ed204425b15c2cf560d086dc

Observation 79c6ad72-af3f-4126-a1aa-0601e6c05547 · outbound

This paper cites G., Guo, Z.

Explicit Preference Optimization: No Need for an Implicit Reward Model G., Guo, Z

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.824269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.824269Z digest=sha256:1347c75aced6692abbbf31e2515285b83147998fd9450c7f8831c6dd3a213ac9

Observation c395d6de-c621-4c82-81ed-498728b3de65 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.828666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.828666Z digest=sha256:440ca38d96def03e4ee3e85d5f7eb53b6aa96bc82075eac4af961c6c5c5253f6

Observation 946f90a8-9e91-4dea-b308-9bf73f94ebfa · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.833973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.833973Z digest=sha256:aa3f31f775a6efbd2dff08beb4316d15426022e283bb7e67032dd8df93aec1db

Observation 99bb216e-2624-4f7b-be33-c84082939d0b · outbound

This paper cites G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M.

Explicit Preference Optimization: No Need for an Implicit Reward Model G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.839067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.839067Z digest=sha256:67e34066d5fd47bddb0b767f296f73dc1a416dcf79ded6950bb73a4710d69d40

Observation c0ef421c-5e0e-41ef-a713-4ec9ce79430c · outbound

This paper cites and Rinaldo, A.

Explicit Preference Optimization: No Need for an Implicit Reward Model and Rinaldo, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.973562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.843370Z digest=sha256:1c1998294e0824db3a5f7837c201c64ba48acab7e2be21a5d16de7181051b537

Observation a615d5e9-8a6f-43da-84cd-91beb628846a · outbound

This paper cites an unresolved cited work.

Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.848212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.848212Z digest=sha256:1bcb993edf6a9d2501b9b37539f745c599e1e471560df0e06687b81ee3bf5924

Observation d0320c2d-dee5-4d5f-b887-0426e3b56aea · outbound

This paper cites T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M.

Explicit Preference Optimization: No Need for an Implicit Reward Model T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.946737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.853062Z digest=sha256:81186a07c39a1a97f00aa17f210d161bf5bbbd43a17ebc4c6f181936af7b59af

Observation f4011480-7ab6-4093-8a94-ac2d90a9758d · outbound

This paper cites A survey on evaluation of large language models.

Explicit Preference Optimization: No Need for an Implicit Reward Model A survey on evaluation of large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.932346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.857050Z digest=sha256:ecd7f1b02d2ffb46bf71c9da14c41c8d3e07d405b6993cbd282b4659e1e6638e

Observation 48ce16f0-6567-4693-b4f1-353b408a368f · outbound

This paper cites Bootstrapping Language Models with DPO Implicit Rewards.

Explicit Preference Optimization: No Need for an Implicit Reward Model Bootstrapping Language Models with DPO Implicit Rewards

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.861184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.861184Z digest=sha256:d2efdf8eee77b09c59f88f98a5a3969ac3271211814d95ec0a3b0967779b55cf

Observation 7bcc5977-6d0d-40de-9a89-38f77c56e04e · outbound

This paper cites Ultrafeedback: Boosting language models with scaled ai feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Ultrafeedback: Boosting language models with scaled ai feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.916344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.865832Z digest=sha256:2d7bf5499b23ca94c990d1a5c552688183dbca3f5e7ff22192ff49fb5cad3b30

Observation 23d7df53-4660-4f6d-8841-aae19e1168d8 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

Explicit Preference Optimization: No Need for an Implicit Reward Model Enhancing chat language models by scaling high-quality instructional conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.869779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.869779Z digest=sha256:dd7ee7992d2b844d545dc3bc15fdb3fd4c673a8179635c5a115df5c863c0fb5e

Observation b3d3c575-fa88-4367-a5b8-4c98693dcc9f · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Explicit Preference Optimization: No Need for an Implicit Reward Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.873960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.873960Z digest=sha256:fa8604072f0df2b6cfaf3018dd7a3ee8f2c508c7eb4bf6ee552a8cede9fad3a5

Observation c7b3bf60-134f-42c1-a96a-de45a050d49f · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model KTO: Model Alignment as Prospect Theoretic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.878001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.878001Z digest=sha256:266e1fccc87da679a7fa54dc6844b9948d8564bfb43bb79e6e5ef62c500950e8

Observation 29c4a37d-1f34-468f-90e3-68cb338f922a · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Explicit Preference Optimization: No Need for an Implicit Reward Model Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.882471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.882471Z digest=sha256:8f5ec1655ff9636f8814beecfac4832c69dffe73fb605e8605258668cbf24979

Observation 9ac79af2-8955-4e80-8e72-19b881aba6d1 · outbound

This paper cites Bias and Fairness in Large Language Models: A Survey.

Explicit Preference Optimization: No Need for an Implicit Reward Model Bias and Fairness in Large Language Models: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.886568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.886568Z digest=sha256:bbfa4e20d199bd4d435971867e3018a1d5f5bd16d38eb53da1ed052cc6127b7f

Observation ee9c6377-7ddb-4e9b-8ca8-df6af4d58ad9 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Explicit Preference Optimization: No Need for an Implicit Reward Model Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.890649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.890649Z digest=sha256:d49fa2b2071aa74a0c43483f6ff517f57351e59f80d4cc0732abdf33b8089413

Observation 9bf454c0-d0c9-470b-99ac-b9b245bc4f18 · outbound

This paper cites Learn Your Reference Model for Real Good Alignment.

Explicit Preference Optimization: No Need for an Implicit Reward Model Learn Your Reference Model for Real Good Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.894671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.894671Z digest=sha256:404d34b7ac2cbbcda721aa82062a37ec37be64e090c457508e59822c5524802a

Observation 186c1af7-8365-4919-945b-3665b82d0bcd · outbound

This paper cites and Pierskalla, W.

Explicit Preference Optimization: No Need for an Implicit Reward Model and Pierskalla, W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.890064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.898762Z digest=sha256:09cc3a612749f4b0cc25ddbb22e3516a3ac3d48c6a45c4a5a387ad9ca7e1b580

Observation 69ecbe00-bcf7-4428-91ab-5fbd69ed5fa4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Explicit Preference Optimization: No Need for an Implicit Reward Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.902738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.902738Z digest=sha256:85409d9215341181369466e243e00ce93ab4a48636b328bdeb40fe522b5e279b

Observation e897ace9-7697-498a-ba60-dd00acfaa9c8 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Explicit Preference Optimization: No Need for an Implicit Reward Model ORPO: Monolithic Preference Optimization without Reference Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.906760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.906760Z digest=sha256:d453a1385e44bd8e3ea4f53fc6b518e30f38b519b753431b0adc7db4259eae47

Observation e1a6ffa6-1c87-438c-9b8e-5c3b6a11a067 · outbound

This paper cites Understanding the Learning Dynamics of Alignment with Human Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Understanding the Learning Dynamics of Alignment with Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.910923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.910923Z digest=sha256:141347132d9b0ce1e1d48f2a38a4fd0166d2d4e8f5e889b2e386075d9dd51eff

Observation d6a10d37-3832-4f5e-b2e4-bf3e7bf18d1e · outbound

This paper cites https://github.com/huggingface/trl/pull/1265.

Explicit Preference Optimization: No Need for an Implicit Reward Model https://github.com/huggingface/trl/pull/1265

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.875095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.915435Z digest=sha256:2171d465e8ce90809bf0c094795d8ff4f3405540eeecdaf9dbd1327f40eafb46

Observation 65a14ff0-d858-4efd-8bd0-57f9600c6629 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Adam: A Method for Stochastic Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.920012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.920012Z digest=sha256:35731091ed97c06b2fa50f0829f2b4bebbfe0700f0feca902c5bd29274701eca

Observation 8dba99e6-d37a-40a8-8f2b-710645e6e6eb · outbound

This paper cites Common learning constraints alter interpretations of direct preference optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Common learning constraints alter interpretations of direct preference optimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.860673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.924252Z digest=sha256:dae971e247bea89b5a8e5bca51eb38fbd5b6ac82914cc471c08f11a05fd58f9e

Observation 06e98cfe-841e-411d-b3be-83c479524cec · outbound

This paper cites H., Gonzalez, J.

Explicit Preference Optimization: No Need for an Implicit Reward Model H., Gonzalez, J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.928495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.928495Z digest=sha256:7f94752807448f18924057ded0972bb29209aa89627d7e584b3d952c25bd35a5

Observation 9daa881e-8c3b-4878-b8ce-b212ac68cc8f · outbound

This paper cites an unresolved cited work.

Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.933007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.933007Z digest=sha256:196dd87c421c72bfa81b252f812f975a6a14394649e1ccf51c3d2c26c5aef51d

Observation b7c6d022-71b5-42c7-90d9-c7b5bcbd23f7 · outbound

This paper cites Policy Optimization in RLHF: The Impact of Out-of-preference Data.

Explicit Preference Optimization: No Need for an Implicit Reward Model Policy Optimization in RLHF: The Impact of Out-of-preference Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.937230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.937230Z digest=sha256:b36bfc9b5d39e26196501d0a5bf2c46c1e778f9d57113366accd3af37ec4bd3b

Observation 08f0ace7-0ead-474d-aaaf-3a7fb6b1e2fe · outbound

This paper cites On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.941137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.941137Z digest=sha256:04165b3e1d87fb8971de9087f82744d17c0e587be0a5e3bc60d10a059eb7267d

Observation 9735fcff-1b7a-4138-b7d9-0ee2d8104416 · outbound

This paper cites L., Daly, R.

Explicit Preference Optimization: No Need for an Implicit Reward Model L., Daly, R

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.825067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.945451Z digest=sha256:f082a8b094849ca08570b987f03fc44a685a8ee4a5fce344d0058a72d133a10c

Observation 3eaad74c-b477-4f61-ad2d-2514ce7af7ab · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Explicit Preference Optimization: No Need for an Implicit Reward Model SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.949816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.949816Z digest=sha256:aa86a1468b6037f10dc82fa07b4dd5f8ca2d87031fe9508625eb53cbee7d85c4

Observation a4d29e7f-976f-40d7-9278-28a683cbe2cb · outbound

This paper cites Active Preference Learning for Large Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model Active Preference Learning for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.953985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.953985Z digest=sha256:2f9efcc9564aeebe7ab3e946dcf55d330ce13fe0bef5450b9e7fd9d18da37591

Observation 31a19b43-e3f6-46ee-8746-81b6e9241408 · outbound

This paper cites Training language models to follow instructions with human feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.958243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.958243Z digest=sha256:3393b831238e0cd4d24d45a2320a89f6e93fb62fed6ffe064fcb93a8a6f2e88c

Observation 59fc338f-5672-41ee-81f7-61e83c74674c · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Explicit Preference Optimization: No Need for an Implicit Reward Model Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.962472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.962472Z digest=sha256:aaf6436c09c494de6229df7e3bacdad2014106fe5bc280603f3c1191d5c66203

Observation 2444c89a-c34f-44c2-a429-af7a17a9eb79 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Disentangling Length from Quality in Direct Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.966403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.966403Z digest=sha256:eb7f02241ed724da17f7d50f27acce93c849a6823d7970b4cb0d0504125225d6

Observation 491e4e82-0774-4f99-87e1-63e182a5feaa · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Explicit Preference Optimization: No Need for an Implicit Reward Model Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.970581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.970581Z digest=sha256:15a1641f451309931323bb4219eeef0fe1350048c150b49ef237263739b8b94e

Observation 8aa11bc4-da3e-4fd0-9bad-4b0f8063e226 · outbound

This paper cites and Schaal, S.

Explicit Preference Optimization: No Need for an Implicit Reward Model and Schaal, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.799960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.974749Z digest=sha256:6a5e98fb43af2e01140bb22415b221d00a6aba283b9b66c41a0d6f503f8a886b

Observation 1b783b59-d2c2-4763-b8b5-40991e1dd7b1 · outbound

This paper cites D., Ermon, S., and Finn, C.

Explicit Preference Optimization: No Need for an Implicit Reward Model D., Ermon, S., and Finn, C

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.978690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.978690Z digest=sha256:c3b252fab633a952f6ae590f76cd516b399586e61e8ab57576869208d2121a06

Observation e4315a4d-6138-4a5e-a77a-56c8f4f8748a · outbound

This paper cites Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.982800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.982800Z digest=sha256:ba19c59c7c0173d6633a70add163be4726ec4742510cb312f407606de22d1727

Observation d7a20d07-a754-48c3-9370-5e3d676483f5 · outbound

This paper cites an unresolved cited work.

Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:40:18.774723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.987548Z digest=sha256:148c59332ae47fc2e0226618579e67ca8b24535727dc53c714f711a87f2cf630

Observation 62dbb63e-f8b3-434a-8ba4-510035797112 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Explicit Preference Optimization: No Need for an Implicit Reward Model Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.991410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.991410Z digest=sha256:dfe3917b20cf013ab7efb60aa1df5e8f6e48b5856b224e55f12b2c580bb82ce1

Observation 9f0c31c3-9de9-486f-a220-7e3369074278 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.996101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.996101Z digest=sha256:25878f89fcf187bcfe7ed59c78399de7fe4dc8976eec40a120afc51e6766feb4

Observation 1e8ed1e2-e9a0-46cb-861c-81de59126bae · outbound

This paper cites The importance of online data: U nderstanding preference fine-tuning via coverage.

Explicit Preference Optimization: No Need for an Implicit Reward Model The importance of online data: U nderstanding preference fine-tuning via coverage

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.760164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:18.000075Z digest=sha256:94567ed3c7e1da942a89d089cc3f59a2db635364cd9e1dcdb78df419e274c5e6

Observation bf476f93-313a-4550-9249-1b0ca218ec18 · outbound

This paper cites M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P.

Explicit Preference Optimization: No Need for an Implicit Reward Model M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.744990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:18.004100Z digest=sha256:39bf7a4f529d8e54e3f8acafb6713f40e871c080a2c3c9515389bb3b368271bc

Observation ab772d2f-269f-40f3-bd2f-4188f6b385f7 · outbound

This paper cites S., and Bagnell, J.

Explicit Preference Optimization: No Need for an Implicit Reward Model S., and Bagnell, J

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.008082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.008082Z digest=sha256:b16d45f6d368b79a94452f35a6027b935ebffa60d12606e8f9a7109b53001f50

Observation c06bc6f8-2862-4405-a46a-971305cb837d · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

Explicit Preference Optimization: No Need for an Implicit Reward Model Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.011912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.011912Z digest=sha256:711b90eeeb9354c87269fc2047fcb0b93cc5bec94fa7a5be762dcee3dff815e5

Observation 6bf9cade-a35f-49df-802a-f4f1d1c79197 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Explicit Preference Optimization: No Need for an Implicit Reward Model Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.016343Z digest=sha256:e900a13da014a8221aaa43e1f30ec1ed52118e651a7c9dace4cbe33b334600e5

Observation 1b95adbf-7431-45c7-ada4-c04e01e142f6 · outbound

This paper cites Beyond reverse KL : Generalizing direct preference optimization with diverse divergence constraints.

Explicit Preference Optimization: No Need for an Implicit Reward Model Beyond reverse KL : Generalizing direct preference optimization with diverse divergence constraints

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.729750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:40:18.020516Z digest=sha256:15e249c0fe32f2a0eae81520755b074e17c5ad0c9c3e4f7f888e616481945a67

Observation 88349501-5533-413c-8a10-33c1c19c004b · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.024969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.024969Z digest=sha256:98306c7728906f9a1ed07b1e5a62a5207ec63ffd2882e931d8aceb0e06290914

Observation a7cab4d7-a511-4c0c-860c-3f7a5215f708 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Explicit Preference Optimization: No Need for an Implicit Reward Model Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.028816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.028816Z digest=sha256:c5cbfafb952d71df3b59d3e894edb95a14b65b8853e447ef21cb4eaa8719e180

Observation 3613ae8d-bd79-4c5b-9de9-0ce589cb6c87 · outbound

This paper cites A Survey of Large Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model A Survey of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.033007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.033007Z digest=sha256:08acff5921e9f2f251c1ec20d9cf52cef5d58d11053f15d8fedb538f2d2d3169

Observation 96a4e02e-511b-44ed-9d04-e4057629f929 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.036861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.036861Z digest=sha256:6bff64b0777ae4e1a406c571f00d037a29b295940c38c6418c4aefbc745a32c2

Observation 05aafb46-278e-4a2c-b53c-2bf326be5b86 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Explicit Preference Optimization: No Need for an Implicit Reward Model Fine-Tuning Language Models from Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.041204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.041204Z digest=sha256:8306d1d60ad5d587b0d293df8511f434ccc3ec71c5ff8d2aa48d003a27a43f4e

Observation bc3da6ce-ceed-4082-a2d3-b4792f7fe476 · outbound

This paper cites write newline.

Explicit Preference Optimization: No Need for an Implicit Reward Model write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.045861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.045861Z digest=sha256:3b4e7d4894a7281ef3e9c33fd7888265507852a20b2db320c7f1932cf98a7d01

Pith citing papers

Observation 8bb23c37-e426-491f-85e8-4e5cf80e523b · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Explicit Preference Optimization: No Need for an Implicit Reward Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:01.282186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:2d31a0c242ef9d88b32732b418b60d9f7a58e8ba332701d467401ac98e51a3a9