Pith. sign in

Paper Citation Record · LEDGER

Learning a Pessimistic Reward Model in RLHF

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 3 inbound Pith citation observations for arXiv:2505.20556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20556 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:07.462011Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:11:46.549150Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation de6fb29f-26b0-426e-92ab-19cb013c0f0c · outbound

This paper cites [2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ ).

Learning a Pessimistic Reward Model in RLHF [2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ )

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:09.146820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:01:07.287615Z digest=sha256:b7f715037d47c097ee2ff1b6206c51f453a7cb2c9c96bd2598aadccfe77c81ff

Observation 36e8ddd2-9c26-4b78-93aa-403379b37e37 · outbound

This paper cites The rejection sampling process is effectively a policy as it takes a prompt as input and stochastically outputs a response.

Learning a Pessimistic Reward Model in RLHF The rejection sampling process is effectively a policy as it takes a prompt as input and stochastically outputs a response

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:08.680829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:01:07.462011Z digest=sha256:5bdb01dabe68b5ceb3f09ad1e2885b7f756ab5583c3bcdae074041f09aad27e4

Observation 85ccfc5b-4c63-4600-9e91-6f20604f8e6a · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Learning a Pessimistic Reward Model in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.764200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.764200Z digest=sha256:a54b2458ac478ecac2fa7a63951b1b79b94c07a50f9a0c6ed0691d58fc8f8076

Observation f499fe4f-e977-411a-88ef-e517db30b68f · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Learning a Pessimistic Reward Model in RLHF Robust Preference Optimization through Reward Model Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.861147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.861147Z digest=sha256:4d4e4ee9039de8cd00349d9cd6a2fc12c93af6eda59d67b1739011b7f9372f10

Observation 1a636902-5b51-4f76-86a8-8078c4290c71 · outbound

This paper cites Mitigating Preference Hacking in Policy Optimization with Pessimism.

Learning a Pessimistic Reward Model in RLHF Mitigating Preference Hacking in Policy Optimization with Pessimism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.976843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.976843Z digest=sha256:91dcf2c1deeeba3ce452532ce0d445e479cdbedd15d0ac4349f3a7a344819c30

Observation 6c726fd7-51c0-4dcb-858c-88becab87c76 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Learning a Pessimistic Reward Model in RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.069235Z digest=sha256:5aba2a63d86aec1815c8fc2f971b2f2dc0d89edcf5e720b18da97f9df4a1dd03

Observation 362722c1-ba46-4730-a3e0-bd411823339a · outbound

This paper cites The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization.

Learning a Pessimistic Reward Model in RLHF The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.147815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.147815Z digest=sha256:df01915690f3611cc3900eb61c6530703a853ca1367b1cdf03c79a121c30dacd

Observation 846f9a8e-399f-4bad-bf6b-ea1336025286 · outbound

This paper cites RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models.

Learning a Pessimistic Reward Model in RLHF RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.235610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.235610Z digest=sha256:5418aab02acfc5e4cbdc7a0aa31dc7a420d4a6396c1926f69a7d8a5c419f96f7

Observation cc13c926-2db6-4395-8878-9f430c3aedc4 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Learning a Pessimistic Reward Model in RLHF Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.329676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.329676Z digest=sha256:6d5e06e8dc225aa4469f73251717074ce28887025664a444af66e29a12c2009c

Observation 10740686-15de-458a-8a08-82ef54fc1c73 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.396920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.396920Z digest=sha256:8d66b07d7756063b6cfd645c8b3a22a04a4f1158d441fdca84d8ab0ae86844ed

Observation 3e72c6b5-9edd-4f59-b902-d1674575148e · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Learning a Pessimistic Reward Model in RLHF RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.490143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.490143Z digest=sha256:920ebb192f3e263c89f65a8a58ba27801307c51b65af54032041c80ac7e6b9dd

Observation a2b21913-dd8a-401e-bf9e-3fe44b655998 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Learning a Pessimistic Reward Model in RLHF WARM: On the Benefits of Weight Averaged Reward Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.649860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.649860Z digest=sha256:c8c645fd6984e80f177aa7afb27fb972831dcf45c7cb2504cf2447d744fca8f9

Observation cf583308-af18-4836-b38e-fa24ff026969 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Learning a Pessimistic Reward Model in RLHF Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.734257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.734257Z digest=sha256:1f34c96bcfe7f76f458fcfdd2fd127266a25da685241807c5d3ad7618da09cc4

Observation dd3eca1a-614b-4892-a228-e5e9e4d1e819 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning a Pessimistic Reward Model in RLHF Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.832817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.832817Z digest=sha256:18300406cde987753d8ebbdebbf32ddf6374c5a5f47c846f700302eecbbe90b5

Observation 752f47ed-e83d-4753-81d5-9c3d5d35c94f · outbound

This paper cites Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation.

Learning a Pessimistic Reward Model in RLHF Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.949991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.949991Z digest=sha256:d80a5d7df861c0fa2143301488676c4d3fd4d9737310d0298ad8b8f9f85787c1

Observation 400d8bdf-8572-4ceb-ba4e-f027564c672b · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Learning a Pessimistic Reward Model in RLHF Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.110662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.110662Z digest=sha256:f8a2f7997839c0eace90d0d601e7518ff30a02860714c9ffcee5fd9d0749d501

Observation 2188078f-3ec4-47cd-8dcd-da45e6a778b8 · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

Learning a Pessimistic Reward Model in RLHF Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.259098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.259098Z digest=sha256:c0a6c88ea75db224a42eb4e183fb8a8cfb1439db7762b2705d0575b048da62b9

Observation efd341ea-140e-4c02-b775-92ebc86fb584 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning a Pessimistic Reward Model in RLHF LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.403374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.403374Z digest=sha256:650b50181f01da2443e0a7dee503d664f7263a8552b9959ad3c3a6fe58d2ce47

Observation 7defe75f-a49b-407f-84a5-9bff0f301c07 · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Learning a Pessimistic Reward Model in RLHF Making RL with Preference-based Feedback Efficient via Randomization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.516230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.516230Z digest=sha256:a45d7c261c0e19b238e801e2f32dda947d8108f5fd6446d2cfe093d3a6c0c3ef

Observation f7e99903-dfd7-4e45-a6ac-6fd3782b22ad · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Learning a Pessimistic Reward Model in RLHF Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.669955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.669955Z digest=sha256:0bb7f01081f1ee34e337951838f3eb018eadacb43d4e43f63ef726670f1e51b5

Observation c925f773-402e-4005-aad0-6c67f4ba00c6 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Learning a Pessimistic Reward Model in RLHF Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.783481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.783481Z digest=sha256:dbc1edba99e5f7dbe68c6de27437ec076018041d7d19d586b992e53d9371f55a

Observation a3074639-539a-49a7-97b8-496c5ef4f467 · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Learning a Pessimistic Reward Model in RLHF Provable Offline Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.936809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.936809Z digest=sha256:253327a8c0847bf5dc4257806243f63fc9ebecae40493dc67916128918209214

Observation 9a3474c4-bb56-4ffb-89d4-bf52cc09386e · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Learning a Pessimistic Reward Model in RLHF Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.063483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.063483Z digest=sha256:6794e6f1c69953b28b154f8d9a9735b3bbfab5da427a1f92c10dc752e58d07c0

Observation 43c8e79b-099b-42df-b0df-f7e4b8b3a09b · outbound

This paper cites Logarithmic regret for online kl-regularized reinforcement learning.

Learning a Pessimistic Reward Model in RLHF Logarithmic regret for online kl-regularized reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.182899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.182899Z digest=sha256:38b627ddf1bbc4faff3bd95d73017d1380d5172d3b6b3ff2b5cb08a69e127399

Observation 192b52ba-3723-4214-ab5e-f95fd3bc5dcf · outbound

This paper cites 1 will increase the bound on the performance gap analysis, which is undesired for sampling efficiency.

Learning a Pessimistic Reward Model in RLHF 1 will increase the bound on the performance gap analysis, which is undesired for sampling efficiency

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:08.907969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:01:07.361030Z digest=sha256:8760c9cb1d6f2dc9cbe0a02ba36ea921e13e1b9a591cd25804b429609abaf247

Observation df8a4b7b-f1ed-41fc-98b0-148f8d11c3fb · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Learning a Pessimistic Reward Model in RLHF Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.563716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.563716Z digest=sha256:c4319ede34e173efbf7d2b27c9879e92ca6663b767fd2a0f5e9a894fb430513f

Observation 3d243c16-4172-4782-96f3-490df15f7716 · outbound

This paper cites Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization.

Learning a Pessimistic Reward Model in RLHF Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.624213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.624213Z digest=sha256:281e3c8f073cda7f142bd8378e178027eef43e7139d5d31e7558b242f68f9a0f

Observation 301e7a58-e6e9-434e-9935-0ccf2266b1bf · outbound

This paper cites Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling.

Learning a Pessimistic Reward Model in RLHF Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:01:08.191682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:01:05.568274Z digest=sha256:4338519ffbead753636c0855746fb4eebf79c7d2459aa9009d75ee5cbbf0f211

Observation 44fed6cc-8e4a-4b34-b7fe-75eaa37a3fbd · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Learning a Pessimistic Reward Model in RLHF Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.472540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.472540Z digest=sha256:45974526bed37db4c6bee90110ad3b9ab50cd72b05e73a16dcbd2a44ef7c96be

Observation 8a6d0d9c-1030-42b3-add9-adf75ffc89b3 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning a Pessimistic Reward Model in RLHF Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.407989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.407989Z digest=sha256:4fbd42dae09c4f0a5896af530aa1d56620ae1a954a0234711915f8650505bf3f

Observation 81bbaf25-9f91-4f8c-a513-b6862e9d21ba · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Learning a Pessimistic Reward Model in RLHF RLHF Workflow: From Reward Modeling to Online RLHF

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.698570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.698570Z digest=sha256:c6dc619a287c005b1c49bd67ac26bf98ab4a7d953011ad143c27f0416e54b609

Pith citing papers

Observation b4b2d56e-a2c4-49c5-af8c-331bc75722da · inbound

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback cites this paper.

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Learning a Pessimistic Reward Model in RLHF

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:46:47.349064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:03:53.045304Z digest=sha256:2f6eb709aa836a437cf3c677ddbe3b6493a67a16e76264eadc2c8ee8a2e6a5af

Observation 3eb09b5e-b6d2-4617-b84c-8e79bd1ac3ba · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Learning a Pessimistic Reward Model in RLHF

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.460204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:b6d4c951a9fc60b40711d47c33abce46449f2c48cea514ace3a15e7642bba005

Observation bfa36826-0ee0-4a04-925f-e4fd9f11aaef · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF Learning a Pessimistic Reward Model in RLHF

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.825574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:f97370abd4f284a266b0c9b7b8889407e3c81874614ce48fb48db1af4c09cde2