Pith. sign in

Paper Citation Record · LEDGER

Explicit Preference Optimization: No Need for an Implicit Reward Model

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.07492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07492 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:18.045861Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:49.149334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:01.276341Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 814c293d-8351-4c84-a266-071dd259de96 · outbound

This paper cites GPT-4 Technical Report.

Explicit Preference Optimization: No Need for an Implicit Reward Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.805167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.805167Z digest=sha256:e9263bd4e6b3700ee55559fe637baeb3bf559f709f725dca6721602771bede1e

Observation 7733375b-fae1-45c0-94d4-4195aa3d2497 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Explicit Preference Optimization: No Need for an Implicit Reward Model Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.810011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.810011Z digest=sha256:a22e1c8a6b4400d29294a187598fc41b80f1d661f9a3dac6ac6443abf0e87515

Observation 28a73e90-003e-4e7b-b6da-7f358d3c3f51 · outbound

This paper cites Llama 3 model card.

Explicit Preference Optimization: No Need for an Implicit Reward Model Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.814798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.814798Z digest=sha256:269dbc59cdf90e9cdbb61c80c77b8e92e684180d4c32bd284a06757858881d19

Observation 587bad29-1ff8-4a50-9e75-684abf84c47f · outbound

This paper cites Direct Preference Optimization with an Offset.

Explicit Preference Optimization: No Need for an Implicit Reward Model Direct Preference Optimization with an Offset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.819727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.819727Z digest=sha256:c38d5ad2d7f354a3f46d90c53299c5182b2e24d6ecaf4358a24ce202955a2a96

Observation 79c6ad72-af3f-4126-a1aa-0601e6c05547 · outbound

This paper cites G., Guo, Z.

Explicit Preference Optimization: No Need for an Implicit Reward Model G., Guo, Z

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.824269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.824269Z digest=sha256:8aaa7b334281742a947ca5bf80ceb423369bfb075a14cf505c82d4a2b0ac8d05

Observation c395d6de-c621-4c82-81ed-498728b3de65 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.828666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.828666Z digest=sha256:a7c723a584df2bdff1ded8118b6ef7575e05d129aaf844e76497301e95caae08

Observation 946f90a8-9e91-4dea-b308-9bf73f94ebfa · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.833973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.833973Z digest=sha256:f88ae0a1b41cc2902850951eac4aa079e64a9393e0b637b92afbc615350f729b

Observation 99bb216e-2624-4f7b-be33-c84082939d0b · outbound

This paper cites G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M.

Explicit Preference Optimization: No Need for an Implicit Reward Model G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.839067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.839067Z digest=sha256:cc194f6e42cd72038f0204a61d1470d2ae60a92ac7d56549008b32133f4edbd0

Observation c0ef421c-5e0e-41ef-a713-4ec9ce79430c · outbound

This paper cites and Rinaldo, A.

Explicit Preference Optimization: No Need for an Implicit Reward Model and Rinaldo, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.973562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.843370Z digest=sha256:0b8086c145dc573e2be8b35e4c5524aaa7ebe103c49623a76b97b38f642f5477

Observation a615d5e9-8a6f-43da-84cd-91beb628846a · outbound

This paper cites an unresolved cited work.

Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.848212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.848212Z digest=sha256:ecb0ef04391e85e7a2432536c28221917cf0b1d5be4c9c7e6fddf58151c8a0a5

Observation d0320c2d-dee5-4d5f-b887-0426e3b56aea · outbound

This paper cites T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M.

Explicit Preference Optimization: No Need for an Implicit Reward Model T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.946737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.853062Z digest=sha256:15a6455b7292cec7076f5fcbc927a1c772e38a2213b2fac5ef9e10c1cc7ec2c0

Observation f4011480-7ab6-4093-8a94-ac2d90a9758d · outbound

This paper cites A survey on evaluation of large language models.

Explicit Preference Optimization: No Need for an Implicit Reward Model A survey on evaluation of large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.932346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.857050Z digest=sha256:044ae94e341329d249746704f7015dd3d776b16be9e3141bfdae624c2f5214fc

Observation 48ce16f0-6567-4693-b4f1-353b408a368f · outbound

This paper cites Bootstrapping Language Models with DPO Implicit Rewards.

Explicit Preference Optimization: No Need for an Implicit Reward Model Bootstrapping Language Models with DPO Implicit Rewards

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.861184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.861184Z digest=sha256:772604ed9c6fe217b9c9e9e5c809734c07c6fded26c09325f91fd1e3495005ff

Observation 7bcc5977-6d0d-40de-9a89-38f77c56e04e · outbound

This paper cites Ultrafeedback: Boosting language models with scaled ai feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Ultrafeedback: Boosting language models with scaled ai feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.916344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.865832Z digest=sha256:6a02bbac102d71e36a719a7f659471a6e51f727019a1dca2609f074e518214aa

Observation 23d7df53-4660-4f6d-8841-aae19e1168d8 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

Explicit Preference Optimization: No Need for an Implicit Reward Model Enhancing chat language models by scaling high-quality instructional conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.869779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.869779Z digest=sha256:d73537fc22bcf635c75a8610e18406ce03d9eb2506831beb848a02eddd5d01e4

Observation b3d3c575-fa88-4367-a5b8-4c98693dcc9f · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Explicit Preference Optimization: No Need for an Implicit Reward Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.873960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.873960Z digest=sha256:5f0c82cff393059d6efbc1b567bea764c6e0cd9df31c8dda5e943da7b135a8e2

Observation c7b3bf60-134f-42c1-a96a-de45a050d49f · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model KTO: Model Alignment as Prospect Theoretic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.878001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.878001Z digest=sha256:3cab8ba6d18305412d28b2e630ec79f8472f9839d4e2400169f915ba174b8209

Observation 29c4a37d-1f34-468f-90e3-68cb338f922a · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Explicit Preference Optimization: No Need for an Implicit Reward Model Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.882471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.882471Z digest=sha256:b8bff7ffd41f2ebe4930c8339b2a26826caa50d7562c457429a350ec173c18f5

Observation 9ac79af2-8955-4e80-8e72-19b881aba6d1 · outbound

This paper cites Bias and Fairness in Large Language Models: A Survey.

Explicit Preference Optimization: No Need for an Implicit Reward Model Bias and Fairness in Large Language Models: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.886568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.886568Z digest=sha256:885762af68cec349ba09de450fddb70fc552d725df6d467fd388c7614c4455a7

Observation ee9c6377-7ddb-4e9b-8ca8-df6af4d58ad9 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Explicit Preference Optimization: No Need for an Implicit Reward Model Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.890649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.890649Z digest=sha256:d1122b5427829b7a4699aaf247f928841cc289920adea5c4a50812e03463d8c9

Observation 9bf454c0-d0c9-470b-99ac-b9b245bc4f18 · outbound

This paper cites Learn Your Reference Model for Real Good Alignment.

Explicit Preference Optimization: No Need for an Implicit Reward Model Learn Your Reference Model for Real Good Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.894671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.894671Z digest=sha256:301ef980324de972bfe08473f1a0ab1896d251f976a508ffba91286d9c4f8ec0

Observation 186c1af7-8365-4919-945b-3665b82d0bcd · outbound

This paper cites and Pierskalla, W.

Explicit Preference Optimization: No Need for an Implicit Reward Model and Pierskalla, W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.890064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.898762Z digest=sha256:16e2f401bcea56db6a2b949e5d6fa9502de460032455a5c2e8c93b858e461122

Observation 69ecbe00-bcf7-4428-91ab-5fbd69ed5fa4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Explicit Preference Optimization: No Need for an Implicit Reward Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.902738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.902738Z digest=sha256:a856f65c20230daddd40146edabbf3b4a71b8796ba1f4f791225e4ce7d7f682c

Observation e897ace9-7697-498a-ba60-dd00acfaa9c8 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Explicit Preference Optimization: No Need for an Implicit Reward Model ORPO: Monolithic Preference Optimization without Reference Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.906760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.906760Z digest=sha256:0e7541fbd88b08312f7c51922867334657da5644857746dab8bed2efb219c14c

Observation e1a6ffa6-1c87-438c-9b8e-5c3b6a11a067 · outbound

This paper cites Understanding the Learning Dynamics of Alignment with Human Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Understanding the Learning Dynamics of Alignment with Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.910923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.910923Z digest=sha256:ee6d727c6c541722c46ec0eb9abb6fa6850ff25421babd4d4c3dc9b6034c90a3

Observation d6a10d37-3832-4f5e-b2e4-bf3e7bf18d1e · outbound

This paper cites https://github.com/huggingface/trl/pull/1265.

Explicit Preference Optimization: No Need for an Implicit Reward Model https://github.com/huggingface/trl/pull/1265

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.875095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.915435Z digest=sha256:98fa2f1d36851a0715c72abc789eeb0f473a540e688fedec3ff88b52ab10f884

Observation 65a14ff0-d858-4efd-8bd0-57f9600c6629 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Adam: A Method for Stochastic Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.920012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.920012Z digest=sha256:1cba47b212dd3b10fd11657044da1e20ba29eb7140f2b179f97e22ccfb7cd579

Observation 8dba99e6-d37a-40a8-8f2b-710645e6e6eb · outbound

This paper cites Common learning constraints alter interpretations of direct preference optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Common learning constraints alter interpretations of direct preference optimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.860673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.924252Z digest=sha256:a0814444b1af398eaecc07fa8a971dc893b5958f2a3c1185fd1f72c46c66de99

Observation 06e98cfe-841e-411d-b3be-83c479524cec · outbound

This paper cites H., Gonzalez, J.

Explicit Preference Optimization: No Need for an Implicit Reward Model H., Gonzalez, J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.928495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.928495Z digest=sha256:492bfdf96dc0713d3bbb41e1ebc7a08adc78453c0487805fc4ca832baecbc859

Observation 9daa881e-8c3b-4878-b8ce-b212ac68cc8f · outbound

This paper cites an unresolved cited work.

Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.933007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.933007Z digest=sha256:ce10a6d9400230b7736f124fa5cf25f352ef453fcfc30ffc864846e9c78ecba5

Observation b7c6d022-71b5-42c7-90d9-c7b5bcbd23f7 · outbound

This paper cites Policy Optimization in RLHF: The Impact of Out-of-preference Data.

Explicit Preference Optimization: No Need for an Implicit Reward Model Policy Optimization in RLHF: The Impact of Out-of-preference Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.937230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.937230Z digest=sha256:0f77ac3e42d95e10ab17254f7b8eddb4af8b43ab2704ba6a60f5bff08d74ce5b

Observation 08f0ace7-0ead-474d-aaaf-3a7fb6b1e2fe · outbound

This paper cites On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.941137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.941137Z digest=sha256:4a173ffdfcbfade0c9bd28b1cee04488a6fd0845b02eb02c21990b8fd25ce7a0

Observation 9735fcff-1b7a-4138-b7d9-0ee2d8104416 · outbound

This paper cites L., Daly, R.

Explicit Preference Optimization: No Need for an Implicit Reward Model L., Daly, R

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.825067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.945451Z digest=sha256:ddec5fca0e114ab28b4b5962e75fce2055f90cc17f6a21804c7a74781e945c5f

Observation 3eaad74c-b477-4f61-ad2d-2514ce7af7ab · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Explicit Preference Optimization: No Need for an Implicit Reward Model SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.949816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.949816Z digest=sha256:5cd6b480e20dbdef0276623a4d476f7b7a910e63a30bb666af3b328d95853df7

Observation a4d29e7f-976f-40d7-9278-28a683cbe2cb · outbound

This paper cites Active Preference Learning for Large Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model Active Preference Learning for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.953985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.953985Z digest=sha256:9fd9c7f9d6dc08a859751d68ed0112de9c16334d211ebd49fecdf128dfaf9c48

Observation 31a19b43-e3f6-46ee-8746-81b6e9241408 · outbound

This paper cites Training language models to follow instructions with human feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.958243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.958243Z digest=sha256:20e3c9333b180ccfa7b4da56e316bc373ad72b20763c9137f9a80307e0e4800f

Observation 59fc338f-5672-41ee-81f7-61e83c74674c · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Explicit Preference Optimization: No Need for an Implicit Reward Model Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.962472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.962472Z digest=sha256:db001950ed80c7582d1d20f7945df4d2720f7612a6a77301a1f0821cb2436330

Observation 2444c89a-c34f-44c2-a429-af7a17a9eb79 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Disentangling Length from Quality in Direct Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.966403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.966403Z digest=sha256:3dc9b3467cf5633130d6d9f676dc0c4179648a7a9f39229934ee4e34d0b07b56

Observation 491e4e82-0774-4f99-87e1-63e182a5feaa · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Explicit Preference Optimization: No Need for an Implicit Reward Model Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.970581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.970581Z digest=sha256:8f982b1c707f40a420a04ba6d5b52ad41c01d3bd880a5155fbde84f429100e93

Observation 8aa11bc4-da3e-4fd0-9bad-4b0f8063e226 · outbound

This paper cites and Schaal, S.

Explicit Preference Optimization: No Need for an Implicit Reward Model and Schaal, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.799960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.974749Z digest=sha256:8fdc4217aa09d9c48d1eb444c95a548c924ffb4aa3091982c28d8a5cbf488349

Observation 1b783b59-d2c2-4763-b8b5-40991e1dd7b1 · outbound

This paper cites D., Ermon, S., and Finn, C.

Explicit Preference Optimization: No Need for an Implicit Reward Model D., Ermon, S., and Finn, C

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.978690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.978690Z digest=sha256:990d80c8477de6bfbb959d3663d81d224fe841ee3be51700d5428c3a4d67d907

Observation e4315a4d-6138-4a5e-a77a-56c8f4f8748a · outbound

This paper cites Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization.

Explicit Preference Optimization: No Need for an Implicit Reward Model Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.982800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.982800Z digest=sha256:f5b71d206899e66b9421307f288d124d9286871435d47597af66f2e75d213c60

Observation d7a20d07-a754-48c3-9370-5e3d676483f5 · outbound

This paper cites an unresolved cited work.

Explicit Preference Optimization: No Need for an Implicit Reward Model Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:40:18.774723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:17.987548Z digest=sha256:83d863ded188e045e32e30ebb95ef0ebf299d2430526f426b0218ea4e10e05f7

Observation 62dbb63e-f8b3-434a-8ba4-510035797112 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Explicit Preference Optimization: No Need for an Implicit Reward Model Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.991410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.991410Z digest=sha256:7fbda6504e2590e2fd5db59bc1b32ea8147e506e7ddf5b2db6c9ba605d89d52f

Observation 9f0c31c3-9de9-486f-a220-7e3369074278 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:17.996101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:17.996101Z digest=sha256:ba8899f972a359fb3ce724aef5a3bbdb73da4478b8f70b81529b68206c3a9918

Observation 1e8ed1e2-e9a0-46cb-861c-81de59126bae · outbound

This paper cites The importance of online data: U nderstanding preference fine-tuning via coverage.

Explicit Preference Optimization: No Need for an Implicit Reward Model The importance of online data: U nderstanding preference fine-tuning via coverage

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.760164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:18.000075Z digest=sha256:5e9e0bbffe0cb8ae68e5f57829a743175ea48fa4fc0749783966829eaf119807

Observation bf476f93-313a-4550-9249-1b0ca218ec18 · outbound

This paper cites M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P.

Explicit Preference Optimization: No Need for an Implicit Reward Model M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.744990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:18.004100Z digest=sha256:a6f5ef641670f7a4970c92858a31fa52733038c3637530a3bd27097028979f68

Observation ab772d2f-269f-40f3-bd2f-4188f6b385f7 · outbound

This paper cites S., and Bagnell, J.

Explicit Preference Optimization: No Need for an Implicit Reward Model S., and Bagnell, J

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.008082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.008082Z digest=sha256:15f051a90dd62efb384d7a511b4cf712abe8ae053d46a3ab9bda3ab1c4998729

Observation c06bc6f8-2862-4405-a46a-971305cb837d · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

Explicit Preference Optimization: No Need for an Implicit Reward Model Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.011912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.011912Z digest=sha256:90923d36ce7aeb1cf066dbc253b771726dfcfb10536219d9e2a7d773a210ae17

Observation 6bf9cade-a35f-49df-802a-f4f1d1c79197 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Explicit Preference Optimization: No Need for an Implicit Reward Model Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.016343Z digest=sha256:c16647ce076840a982ded85b565dce46723347016ed53df70504cd845da267dd

Observation 1b95adbf-7431-45c7-ada4-c04e01e142f6 · outbound

This paper cites Beyond reverse KL : Generalizing direct preference optimization with diverse divergence constraints.

Explicit Preference Optimization: No Need for an Implicit Reward Model Beyond reverse KL : Generalizing direct preference optimization with diverse divergence constraints

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:18.729750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T05:40:18.020516Z digest=sha256:793ba26b226d85e2ccac2ef9e6089f4f520b5c1aecce32513c0a0b627494d133

Observation 88349501-5533-413c-8a10-33c1c19c004b · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.024969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.024969Z digest=sha256:9d087c1a3ae02281eab6805b1d77b5b6a507c4047c2c7fc4f09b4dd934d888b8

Observation a7cab4d7-a511-4c0c-860c-3f7a5215f708 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Explicit Preference Optimization: No Need for an Implicit Reward Model Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.028816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.028816Z digest=sha256:8b3f2653e6b1237cb18bcb0084816cdc76b79f51654d53a8dd96b6c6fa6c6c3d

Observation 3613ae8d-bd79-4c5b-9de9-0ce589cb6c87 · outbound

This paper cites A Survey of Large Language Models.

Explicit Preference Optimization: No Need for an Implicit Reward Model A Survey of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.033007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.033007Z digest=sha256:c1f2101a7a32820577c81663c3d51eec2cc70d18ed854fa723cceabd044af7d9

Observation 96a4e02e-511b-44ed-9d04-e4057629f929 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Explicit Preference Optimization: No Need for an Implicit Reward Model SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.036861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.036861Z digest=sha256:4d145260c673751afffe5baefeef3d5116769e4b7fcfe2694ef6b95dfcf8508c

Observation 05aafb46-278e-4a2c-b53c-2bf326be5b86 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Explicit Preference Optimization: No Need for an Implicit Reward Model Fine-Tuning Language Models from Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.041204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.041204Z digest=sha256:140db2b1e32447c58a90ec5b192d860a3a26d3ac93c2404399114d11ecb9ada3

Observation bc3da6ce-ceed-4082-a2d3-b4792f7fe476 · outbound

This paper cites write newline.

Explicit Preference Optimization: No Need for an Implicit Reward Model write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.045861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.045861Z digest=sha256:28a2cb18b677bbb8b5bfb95dcf2db2863af193b6c554af213483f4beeea7fa10

Pith citing papers

Observation 8bb23c37-e426-491f-85e8-4e5cf80e523b · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Explicit Preference Optimization: No Need for an Implicit Reward Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:01.282186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:77ed36c4b3d64d46cea2ac7689751b4a79a6ab61d4a5e84b255860be8c03ff40