Pith. sign in

Paper Citation Record · LEDGER

Secrets of RLHF in Large Language Models Part II: Reward Modeling

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2401.06080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06080 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.351960Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73fa1a52-2f38-4c64-a58c-de79b310f68f · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.707958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:598549363d2a0c77f0fc501397d84ae5ca5fc0478febdf8f1dca5d904072b245

Observation 0b09bcf4-e9eb-4ddc-840d-b79250750aaf · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.483570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:e7b73eba6a16ff544838fdafca7d7a6b36cf02e2cf9937b18706715299e35b9e

Observation 6fc7a908-e2fc-48ce-bc07-3bf1539d8369 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.899604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:7721c0ebc83ef5c2f7fd50a6ae2d09d10a4543dc5257629364920e6db8f3d05c

Observation 98c2e227-9c84-4ac3-b7b7-dc299e2ecb9b · inbound

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking cites this paper.

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.351960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.351960Z digest=sha256:77416854914312fd1b0b0ac2f781d4cb53e9dc1d8b463e355c318bbac5c839e5

Observation 791982b8-e588-4887-bc77-7e90d866e91f · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.634354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.634354Z digest=sha256:776b3a5458b8bfd7ed00de960ec732c8b119083e37ba7d6e4ee3c28c1d662f5b

Observation e5842996-e851-443c-bc8d-630bf3aa9e5c · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.959132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.959132Z digest=sha256:32d2a54e7577c7f53226178728228b1e5bc214ef9d871584e017923a018e2be8

Observation 95f05551-4bd2-4d8b-8c57-24adc9473adc · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.575352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:4cb7e1c2e3ac67bf01f5d13edd18ecf8f60f4f4a47b085fc8c64376d59cecd57

Observation 67d604db-332f-4d2e-a985-bce45b67d65c · inbound

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions cites this paper.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.445137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.445137Z digest=sha256:6382f8e8eb1c056aa536fa6a6f6597a20c4b393cd01626bf5027dbbec4a99d2a

Observation b94a2143-d519-4a37-9fa7-fbb3ac0946b1 · inbound

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples cites this paper.

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T11:57:11.983807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:57:11.983807Z digest=sha256:530ecc60a54678d49768453c6916707337a2d3e7673579caefbf39513023069d

Observation c93a8b98-a211-439d-86d1-bf05ae4a3959 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.576663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1d4deaf5b22759f23fc546617ac41e0d28ebe4156f18b041ed9bfd53f6707bac

Observation a812a115-e4e8-4789-8eca-a48adb5d8948 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.624148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.624148Z digest=sha256:cbff179a95778579134fd3c9b3c7e65fef046e557394a5a99ed9b9e83180d961

Observation 987cc69d-84de-4dab-9a7e-544ac5e73b2c · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:31.942943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:31.942943Z digest=sha256:215885510bb466805f962cf3b994182ef4145d28a0f3a74f88a9284e30332cba

Observation 5fea45f8-50fa-4f09-9407-96a73b02ce3a · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.352622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:a9ccb51ce28fb9b9031e398cbcdd5e036c6c1a88013f132179bedd7d1d20fe80

Observation 3952763d-ace7-4824-9572-3163cc78be12 · inbound

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing cites this paper.

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:55.913770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:15:55.913770Z digest=sha256:fdecda400275886c9f4cf4ba6f723089336a9d26beefc6383a7efcb91d87f102

Observation 712e93bf-a4be-4b77-8f3b-32299be42706 · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.965198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.965198Z digest=sha256:c085fda3ded710ff4430972e487e3c191bc680834731eef359dd30d9c9b12cfc

Observation 39dc39f2-8afa-497d-ba50-1883950b60d0 · inbound

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding cites this paper.

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:28.485425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:28.485425Z digest=sha256:e5b173348efeab581f62ed1ad7345e08257f8adeb174ddc266f6ecd40c163c9f

Observation 27fdfeb5-0952-486b-a2b7-02ddb351d597 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.462615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.462615Z digest=sha256:af92e9097ede61933d2ca8f7c12b48441c3ac2249cb470f51823610df09578b5

Observation 6e2155e5-e488-4033-a71f-79949deb1a4c · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:22.806003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:22.806003Z digest=sha256:9a2375ac364b10bd94481d606df8fc70ce8ec5271f4370086e9bb11e878308da

Observation b1810d4c-e22a-4a9e-b1f4-565668a357f8 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:07.012995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:ed221ed99bdcb1568375aae6d5bfd25e1c8b8d73b74c7e8c5aa9c4dcbf8978af

Observation b3c752b3-7b51-4c17-af23-0b4694293918 · inbound

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration cites this paper.

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.419338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:07:20.180349Z digest=sha256:acde811899916b483196fe1597783efa9af1f6177a40c5ce5f86da5b75a3904f

Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.159008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:28bf22b270a0111a3bfa2a438d1dbde3be689c6817fcf4f66fead5e55c05d0b1

Observation d3ebbd09-47a7-412e-97bf-eac652a28814 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.604855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:f1a2c99d2179b420ed8c86b2e4bf00b7852b82618ed16ab022edf01eaeeac331

Observation 8bdc2aa2-39d8-4cc2-a6a1-fa90e214838c · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.869033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:9d653e927c6a897cc5a871943fb85c1c923561ddae97c3726521710cb0ed8683

Observation 3a406e49-0558-4235-92b8-ab3799e9c7d8 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.718303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:11b7ffbb0915f407a93076ca3800de5bc819840b03b957a28b94c4eeb4995d76

Observation 9dcd4027-37e4-456c-b1e5-32c00cd4bbd3 · inbound

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity cites this paper.

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:16:18.568531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:16:09.213612Z digest=sha256:e59995dc8ed406d1bcba3e1494d518cc30a9b1da6ce1396affbfda964241f68b

Observation 9e9cfa9b-6513-47ce-a9ae-ac8768437ab2 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.945673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:ff55956907ce834f9fc95165539c4c19259423f1552a8b668d3091c1610ab10d

Observation 90b60aec-fea5-48c8-a609-4e07d988c717 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.017450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:ce81182b9366410173bfb0c12e40849f2cf0c156eb49b1b2281d7ec85ca91244

Observation 8fe3cacf-d709-4e12-97df-291176f3fe31 · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.296018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:28e66a781d67177b20a23ba86dffd8026be4fd4d556555f1b26b287f153775e4

Observation 7b24150e-32d1-49a0-aa08-65e809dabaa6 · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.957727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:c9d800fe74994145af1b41a92d34a4a3a17baa747c7b28f5a387f69bc7e2f9e7

Observation bcf61e05-b703-4838-96f2-ca6ee50f3ef2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.515245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:fa4e1b5f64240aa68ed95b1312bc0f763223967489c97c18d1733c0f63b3027a

Observation 24da175c-92c6-47ff-9ec0-a4ccd107e1a7 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.282385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:a68a5c1951c017b1b5bf17214ccbe7647fd556f7263ebecf650543102e0c75ba

Observation 84b2b752-56a8-4a70-ac76-0f528db5dc7a · inbound

Understanding helpfulness and harmless tension in reward models cites this paper.

Understanding helpfulness and harmless tension in reward models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:22.991762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:06:23.680572Z digest=sha256:ac28de4813e0063182928b9a70b2734735260738b35461db1f2a6141d2a75525

Observation 33796558-f5e7-46a1-87c9-9ca4b3415e13 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.105153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:4015c79833e3727b4869d189d8d4df8110724ff04d74d43470b03221085f4066

Observation 90eb243c-e305-431b-99ab-b8b2084915af · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.135832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:8555d28165a800e0f2a556ae2265bf296fb9c7948b9281aeaed410013e0cf34d

Observation e39d2557-3329-4f19-bc1b-9b02bb77a51b · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:b11301c696783b8b71051bb2bdacb771d36ee8b367e6808d4539a6519da46a33

Observation e433b4e7-168e-4078-ad68-e57a0bc95423 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.827430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.827430Z digest=sha256:62a7da074fa50d78a295f8dacee72b187a3f4435633244087b4b2342de4e22cf

Observation 36eec2bb-fe64-48d6-9715-b767de896496 · inbound

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cites this paper.

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:14:12.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:14:12.745471Z digest=sha256:c35874949ccf1a51a1d450519ba1794788f8d40e86203fef6357b1a08efbb90d