Pith. sign in

Paper Citation Record · LEDGER

Robust Preference Optimization through Reward Model Distillation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2405.19316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19316 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:41.999377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:10:10.421442Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 89e4aabb-baae-45b5-af0f-9e2d9e9d6917 · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL Robust Preference Optimization through Reward Model Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.999377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.999377Z digest=sha256:4c09415948321fc6d00340eaa8f3fa43076b2b1e7b5c2269aabdb92ae3188748

Observation 797bfae3-4e3f-4202-ba83-994136bb68fa · inbound

Frictional Agent Alignment Framework: Slow Down and Don't Break Things cites this paper.

Frictional Agent Alignment Framework: Slow Down and Don't Break Things Robust Preference Optimization through Reward Model Distillation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:00.991880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:00.991880Z digest=sha256:022da73254f698a3f67043f18430e1226316fd03d82de8d3cadb411960dd77e6

Observation 1acd4098-80ed-4623-bbf1-898a97b2c34f · inbound

Risk-aware Direct Preference Optimization under Nested Risk Measure cites this paper.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust Preference Optimization through Reward Model Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.345978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.345978Z digest=sha256:f326d1d64c1a17186831ff9466c0d378adc6e376b3fb8a86b759d39a9b71eb49

Observation f499fe4f-e977-411a-88ef-e517db30b68f · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Robust Preference Optimization through Reward Model Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.861147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.861147Z digest=sha256:4d4e4ee9039de8cd00349d9cd6a2fc12c93af6eda59d67b1739011b7f9372f10

Observation 05043c1f-ced1-415c-9229-4faccb5a5452 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Robust Preference Optimization through Reward Model Distillation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:54.536018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:54.536018Z digest=sha256:de33e878a75fb749c6ff3bc4e33ea9e470299904fbe50fd25d28e476e104b901

Observation 8d036d27-18f1-40b4-9b72-d2a38a1bdf8a · inbound

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities cites this paper.

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities Robust Preference Optimization through Reward Model Distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:25.113671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:25.113671Z digest=sha256:758c873c206763a831ca084bec1ac326661fa4e566bd7ea469c922a60a022b0c

Observation 803e3cde-bdbf-4e6e-849f-6e2179ecd459 · inbound

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion cites this paper.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Robust Preference Optimization through Reward Model Distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:13.397543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:13.397543Z digest=sha256:abd8684b5355c41fb929f1de5a008da67a854bdff954a62ad3a919db7a616b04

Observation 8134630d-ed0f-48c1-b116-42f1f1a20d78 · inbound

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech cites this paper.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Preference Optimization through Reward Model Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.255346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.255346Z digest=sha256:0a40c2d09b1f6eacf0d5f1a401fac296d35447476f9e9a3df482309c5db2e663

Observation 76f6a24b-e98e-4788-ac89-2b54f12811da · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.953881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:c352289cfdd8876c60cb187d9fd121f3c05b99ef34eb309c73419b1cf3653551

Observation ee0af6d8-2c8e-43f7-9456-befc13bb6f26 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Robust Preference Optimization through Reward Model Distillation

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:31:24.845144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:a5d1e829b3d792b6a1933a4e4332d7a753ec300f4d6760a2474da687d7e81601

Observation d8754ce3-7ed1-4c12-b701-f451f8677142 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Robust Preference Optimization through Reward Model Distillation

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:30.780009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:30.780009Z digest=sha256:ba5e745a5c3b269408b82233bb94f5f6feb7ee2260f54b736970f254a213a09c

Observation 7e62ac6f-5cce-443d-a546-a63ae0fd0e8c · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.585241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:14d2dfa1976f5e818ea61e65fe16ad09ceb682ff109c0f9438d4d3dcb0e89209

Observation 76156812-159e-4015-bb05-c44fe90c4f7e · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.423914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:5ed0a74872fa2d4d8401dd5add3eb7c40359481fbba4ebe6ba7ac30979c99b68

Observation 8da4c509-4e9c-464b-b81b-2995af2b488f · inbound

Generating Place-Based Compromises Between Two Points of View cites this paper.

Generating Place-Based Compromises Between Two Points of View Robust Preference Optimization through Reward Model Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.122390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T03:36:31.695964Z digest=sha256:54f9b52c674109b074a537290ff0ad5d62d232bef7d86c76907f8d74024bf34b

Observation 52382d59-bd63-41ce-87da-ffbec13e15e0 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:57.377453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:57.377453Z digest=sha256:b86a8a40034750ed309b3536e41be73997de4287d55bdf0593a3a1d8b0b0d449