Pith. sign in

Paper Citation Record · LEDGER

Robust Preference Optimization through Reward Model Distillation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.19316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19316 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:56:28.717046Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:10:10.421442Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3f408ca-681e-4cd4-8491-8de269854632 · inbound

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment cites this paper.

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment Robust Preference Optimization through Reward Model Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T19:56:28.717046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:56:28.717046Z digest=sha256:cb44fae19cc4efe4ae722eca742e7ed99cc02c01354b08a1281e80093f9c45ac

Observation 89e4aabb-baae-45b5-af0f-9e2d9e9d6917 · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL Robust Preference Optimization through Reward Model Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.999377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.999377Z digest=sha256:8555b7138c95e3475cc01ab413e3766e6416b406c59b43b3ab6689947e524bf2

Observation 797bfae3-4e3f-4202-ba83-994136bb68fa · inbound

Frictional Agent Alignment Framework: Slow Down and Don't Break Things cites this paper.

Frictional Agent Alignment Framework: Slow Down and Don't Break Things Robust Preference Optimization through Reward Model Distillation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:00.991880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:00.991880Z digest=sha256:01067f93d3d965e11a820f7eef1f04ad11f632ebe0001254bf28085322e73b94

Observation 1acd4098-80ed-4623-bbf1-898a97b2c34f · inbound

Risk-aware Direct Preference Optimization under Nested Risk Measure cites this paper.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust Preference Optimization through Reward Model Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.345978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.345978Z digest=sha256:eb954b94d9debabd0b91b1bb239cf1ef9c319b34023c14d21b8781122b185007

Observation f499fe4f-e977-411a-88ef-e517db30b68f · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Robust Preference Optimization through Reward Model Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.861147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.861147Z digest=sha256:4d4e4ee9039de8cd00349d9cd6a2fc12c93af6eda59d67b1739011b7f9372f10

Observation 05043c1f-ced1-415c-9229-4faccb5a5452 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Robust Preference Optimization through Reward Model Distillation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:54.536018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:54.536018Z digest=sha256:e957a92357d12ac1fe5f6a19c3e159870d1f46d760a286ae3e09bfb82ffd8de4

Observation 8d036d27-18f1-40b4-9b72-d2a38a1bdf8a · inbound

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities cites this paper.

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities Robust Preference Optimization through Reward Model Distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:25.113671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:25.113671Z digest=sha256:758c873c206763a831ca084bec1ac326661fa4e566bd7ea469c922a60a022b0c

Observation 803e3cde-bdbf-4e6e-849f-6e2179ecd459 · inbound

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion cites this paper.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Robust Preference Optimization through Reward Model Distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:13.397543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:13.397543Z digest=sha256:75d3c80ba061b799a35292051c1d40560d7411d455cb0d11b12de1863765a74c

Observation 8134630d-ed0f-48c1-b116-42f1f1a20d78 · inbound

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech cites this paper.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Preference Optimization through Reward Model Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.255346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.255346Z digest=sha256:0a40c2d09b1f6eacf0d5f1a401fac296d35447476f9e9a3df482309c5db2e663

Observation 76f6a24b-e98e-4788-ac89-2b54f12811da · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.953881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:e8da28f5b98f900b03b72a7edb862f5e4c226cac5eededa38d53273cc79a1e46

Observation ee0af6d8-2c8e-43f7-9456-befc13bb6f26 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Robust Preference Optimization through Reward Model Distillation

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:31:24.845144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:35b926f76cf4c42025cef1987ad5fa016bfde036ce6532757381d7592866d10c

Observation d8754ce3-7ed1-4c12-b701-f451f8677142 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Robust Preference Optimization through Reward Model Distillation

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:30.780009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:30.780009Z digest=sha256:ba5e745a5c3b269408b82233bb94f5f6feb7ee2260f54b736970f254a213a09c

Observation 7e62ac6f-5cce-443d-a546-a63ae0fd0e8c · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.585241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:02aa603d7a90e74792a8855de77fb549afeb46ff88ba6128546b20b638b72f68

Observation 76156812-159e-4015-bb05-c44fe90c4f7e · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.423914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:35bb761328daf03309429e01188406f40854edfbb3ae2559172968421f1a2f48

Observation 8da4c509-4e9c-464b-b81b-2995af2b488f · inbound

Generating Place-Based Compromises Between Two Points of View cites this paper.

Generating Place-Based Compromises Between Two Points of View Robust Preference Optimization through Reward Model Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.122390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T03:36:31.695964Z digest=sha256:4f9c495820a4b20cab6be6b23ad12d834036e5b8ca51430f591ff939eb5f19bb

Observation 52382d59-bd63-41ce-87da-ffbec13e15e0 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:57.377453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:57.377453Z digest=sha256:b86a8a40034750ed309b3536e41be73997de4287d55bdf0593a3a1d8b0b0d449