Pith. sign in

Paper Citation Record · LEDGER

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2504.04950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04950 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:54.497908Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:46:14.356541Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6b797002-0f2c-4569-b19c-7217b5948dae · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.751164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e65099d1bf7e0cbf69d526bc0f3f3e609b9ec26d2809787836be48a3d22e95b7

Observation 79be537c-ece5-4495-b294-07b595d9fd67 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:54.497908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:54.497908Z digest=sha256:910d429043147a39a0057b202a72a536cb38d0133095697602556f3c5e735255

Observation 8e894e89-aaa8-40d4-a55e-a065552fb817 · inbound

WebDancer: Towards Autonomous Information Seeking Agency cites this paper.

WebDancer: Towards Autonomous Information Seeking Agency A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:00.584149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:00.584149Z digest=sha256:2309111ee1d392031715e1080e651b3d02e9bbc33150f65483d9cc3f859c5288

Observation 972cb59f-2d79-474c-be90-22b1351106c5 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:37.841550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:37.841550Z digest=sha256:6a8b36cf4b41c0a85c536a9dffb5a61d8912ee74d24e9a19bbd22b81e4dd0301

Observation 96ff8a98-aaa1-4426-84c9-7d72a24e8711 · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:09:01.240055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:09:01.240055Z digest=sha256:0b58ff197cc64b125ff27c0493ed49b963448b78f3627ab2b10ecae18e4524cc

Observation 64bc43cc-a96f-4ce5-b2fb-14d01e7ae83d · inbound

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization cites this paper.

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T09:28:13.017326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:28:13.017326Z digest=sha256:950d87c555561a337217ea141c05a0b5e599c04c8e62ea01521f93e49251d7e7

Observation aa743bd7-6033-40e5-8722-9cbe0250b7bd · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:20:49.249595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:01e056fe382533fbcff6bdbdd20bc1d9b2978d56548b23afea87f68f61a5dff4

Observation f7d6ff37-2f61-492e-a2ec-ca2b96a8755e · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.558499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:6b14a1eea60543cae44ee9eea8a2446f88b848a63292d8d5b3e3720b4becd734

Observation b758234d-2223-4943-955d-1e58f95674e2 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.945999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:3603634e1776148977ac714fae30ec82caf0f909b2565e750f8bd9e2708dcc74

Observation 16782ab1-f999-4139-bbfe-0df663bba745 · inbound

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation cites this paper.

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.891575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T10:15:22.633882Z digest=sha256:559918f4d88aced2cda8ea76434a9510a6d4886e2ab48a2ae24a25b1328cd11a

Observation 5104f631-dd80-4742-9611-63305d747dbf · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.358178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:b238c9537839bd3199753ac5ee54d2c079e90c97698d15aaf9baf0638b7dfc27