Pith. sign in

Paper Citation Record · LEDGER

Secrets of RLHF in Large Language Models Part II: Reward Modeling

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2401.06080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06080 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:35.445137Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73fa1a52-2f38-4c64-a58c-de79b310f68f · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.707958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:b55fdcb7730911c426e352cbeecd6c4f8c413568c164db8afa5b9e69cbaac9a5

Observation 0b09bcf4-e9eb-4ddc-840d-b79250750aaf · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.483570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:633165d3412c258fb0ae720aa40b65249ec7f76d3605c3da4e54918145d3b724

Observation 6fc7a908-e2fc-48ce-bc07-3bf1539d8369 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.899604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:5d89a3fbf4fb8a4a0afec5a4ab11eedec99ef565f61c60cc1998b64646e9b922

Observation 95f05551-4bd2-4d8b-8c57-24adc9473adc · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.575352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:094dbabb81f5a7f51b661d28104f601ff0c97006c603119f2c26d8cfaf07f251

Observation 67d604db-332f-4d2e-a985-bce45b67d65c · inbound

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions cites this paper.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.445137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.445137Z digest=sha256:59f425a119acfd2868e3093353c26703e4f2d4cd15044a4edb9b6c6686ff7316

Observation b94a2143-d519-4a37-9fa7-fbb3ac0946b1 · inbound

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples cites this paper.

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T11:57:11.983807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:57:11.983807Z digest=sha256:17feea099fe8ce472eb8acf32534785264a1a9688f0720f1b5e2f7a05c6297f5

Observation c93a8b98-a211-439d-86d1-bf05ae4a3959 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.576663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:10beb17ce3e590a98d122636d7f528423068bad5f7a25bd70c337bdfcb8663ac

Observation a812a115-e4e8-4789-8eca-a48adb5d8948 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.624148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.624148Z digest=sha256:827af363072c78d53069c89f3dd38ca0918f452f534da798b4e938d7ea7bb776

Observation 987cc69d-84de-4dab-9a7e-544ac5e73b2c · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:31.942943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:31.942943Z digest=sha256:acbebb2fb018200b4b58d9d8a4b652470d850a8386c8a62a1b34f951cf3668bd

Observation 5fea45f8-50fa-4f09-9407-96a73b02ce3a · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.352622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:b859a493e3d2ed853d13a36d525db06d430cbe553817605127c100e3e95b140d

Observation 3952763d-ace7-4824-9572-3163cc78be12 · inbound

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing cites this paper.

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:55.913770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:15:55.913770Z digest=sha256:fd1d51de1ae41059e05ba2f9a0a5a533c4b6c9f2a8484766b17514e5d9919294

Observation 712e93bf-a4be-4b77-8f3b-32299be42706 · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.965198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.965198Z digest=sha256:1b2a3c5ad0699f22665ceadfb2c674393e572f0e3b61df47af44ada03c153bc0

Observation 39dc39f2-8afa-497d-ba50-1883950b60d0 · inbound

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding cites this paper.

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:28.485425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:28.485425Z digest=sha256:2826e3755e5922ffe0d8b10c4b497d9f131fdb6034ff90cd47f201ed08374297

Observation 27fdfeb5-0952-486b-a2b7-02ddb351d597 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.462615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.462615Z digest=sha256:c8773e3cc2367c55d2462f9057480ad2d63600fc66465f14e971a67dee0500f4

Observation 6e2155e5-e488-4033-a71f-79949deb1a4c · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:22.806003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:22.806003Z digest=sha256:a98a7c3a6df96899f541ce7f610a541984a4bb249d378d1edc49a315ab1b7ae9

Observation b1810d4c-e22a-4a9e-b1f4-565668a357f8 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:07.012995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:4e8eeabc37000c0545c59ffb622b8cb6d900e2f13e907afbec4051de60ab35e8

Observation b3c752b3-7b51-4c17-af23-0b4694293918 · inbound

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration cites this paper.

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.419338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:07:20.180349Z digest=sha256:a671ee767e66c1d817a063715ea6cf9949c4728df18d8617552971d432264e4c

Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.159008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:fdc15e192c897b4e29e9fc5636a6a04c9746e98a236693e0b1f30eb559547955

Observation d3ebbd09-47a7-412e-97bf-eac652a28814 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.604855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:a926bafae31f0f1f19daa0e1ad0159647b3f6ede54dcee04e0c39ed176574fe6

Observation 8bdc2aa2-39d8-4cc2-a6a1-fa90e214838c · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.869033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:4a4aa4c68b9524d2a429da1491a36a8c35b11ee940b348465a2e476942d3958f

Observation 3a406e49-0558-4235-92b8-ab3799e9c7d8 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.718303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:33e61b70005c750ea3685f2bb7801c01c05229b65a10c30543f5be5ae3e74aa7

Observation 9dcd4027-37e4-456c-b1e5-32c00cd4bbd3 · inbound

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity cites this paper.

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:16:18.568531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:16:09.213612Z digest=sha256:7c60d529ca5e6b51dd0501e652f4a673285dde8fb8332622aea74d960f2ae587

Observation 9e9cfa9b-6513-47ce-a9ae-ac8768437ab2 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.945673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:e367e4307e8e323231d8ccee3cef384de77280a19267136503abf2a318ef3be3

Observation 90b60aec-fea5-48c8-a609-4e07d988c717 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.017450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:735c394f279f406494da689aa606ddcb09b6f0231d841159b6f960f661f93566

Observation 8fe3cacf-d709-4e12-97df-291176f3fe31 · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.296018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:bece9135479f44ed7fc1f150d8b27648f5235def036bfac92900de9bacae632b

Observation 7b24150e-32d1-49a0-aa08-65e809dabaa6 · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.957727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:10d16792f48fce1e9c3a262be8fe605962b10ef686e4e3b0e55ccdfba4fa1fde

Observation bcf61e05-b703-4838-96f2-ca6ee50f3ef2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.515245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:f6dcd02834d7b28f5b507c6844e1c4fb30e89c5a519ee2aaa394082cf938d10e

Observation 24da175c-92c6-47ff-9ec0-a4ccd107e1a7 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.282385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:9652fb266c19b1190d9b827cdaef28b212783920d7510d1cfb98d4092b8d8711

Observation 84b2b752-56a8-4a70-ac76-0f528db5dc7a · inbound

Understanding helpfulness and harmless tension in reward models cites this paper.

Understanding helpfulness and harmless tension in reward models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:22.991762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:06:23.680572Z digest=sha256:09b4d03b6357109b27e45d8be81da0cd0c0a3cf21736838ce3db463eab4c38ae

Observation 33796558-f5e7-46a1-87c9-9ca4b3415e13 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.105153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:f7c336f9bdd460c44f014b505382297fbded6f49a6a25be4ed640c4df9b9b1ef

Observation 90eb243c-e305-431b-99ab-b8b2084915af · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.135832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:cf5c2fbdb0d020edd9ccd562d7305a35c417b01e8a7dcbdcfb543d4a82d590d0

Observation e39d2557-3329-4f19-bc1b-9b02bb77a51b · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:0525599fe7f9b77cc6eaeb5660a537392f5e06b5b80d814042c6d06f9b27e7cb

Observation e433b4e7-168e-4078-ad68-e57a0bc95423 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.827430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.827430Z digest=sha256:cb21ce151690e90bc936e1556b667a31201a356af2b73b2ccbbfd109cc991724

Observation 36eec2bb-fe64-48d6-9715-b767de896496 · inbound

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cites this paper.

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:14:12.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:14:12.745471Z digest=sha256:2f6bd10d9408a20258f830640dd5b5b1509af2a386bd77c36b1e893454aa8b73