Pith. sign in

Paper Citation Record · LEDGER

Accelerating RLHF Training with Reward Variance Increase

As of 15 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2505.23247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23247 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:49.168492Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77bbfb63-640e-41ac-880e-8354745ec0f1 · outbound

This paper cites GPT-4 Technical Report.

Accelerating RLHF Training with Reward Variance Increase GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.286507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.286507Z digest=sha256:9b91e86e9bece6492c6b03ee191cc94daeaebf797d052aeb6a86747c7f20f670

Observation 5cd1c3b0-8545-40f8-9e76-ab1e7b05cc95 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Accelerating RLHF Training with Reward Variance Increase Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.338823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.338823Z digest=sha256:a85d4ae819725fa4d1b86ed259a1c7766076877001df85f16ee0017f46742b0b

Observation 5df03760-eb3a-4895-bc6e-e7617d7c849e · outbound

This paper cites Biderman, H.

Accelerating RLHF Training with Reward Variance Increase Biderman, H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.930185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:44.398065Z digest=sha256:d164c3750f2f6537d0f03f27cd690dcf14b55d7501005a18e3d25db9fd494292

Observation 009e307f-77ea-4bc4-b906-2638e8a9c2b4 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Accelerating RLHF Training with Reward Variance Increase On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.493191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.493191Z digest=sha256:93075ddb41f532e55906acf2671af57d8aec863930e2a6b81377c78dbb917233

Observation 873fd355-ecdc-4e96-bc55-42f8314b04fd · outbound

This paper cites Brown, B.

Accelerating RLHF Training with Reward Variance Increase Brown, B

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.702380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:44.558034Z digest=sha256:6d8bafd919b84c1a7f34ad7291b56cba08d74130de092626340273b931f35238

Observation 478a1110-21f8-48ec-80f5-02d5da0601ff · outbound

This paper cites Busa-Fekete, B.

Accelerating RLHF Training with Reward Variance Increase Busa-Fekete, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.492367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:44.669918Z digest=sha256:01bb9f7f440a7846893fd9e3101795247f5ffd97e041816d836c61d5e1f630eb

Observation c28f5f59-6cf7-43b9-96d3-637382b39030 · outbound

This paper cites The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models.

Accelerating RLHF Training with Reward Variance Increase The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:50.084647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:44.776384Z digest=sha256:2ad3994d607deaf27e1ce0b3019abde4273f3a9e6314f335dec5f2aebb27b309

Observation 54bf522b-4f9c-4680-a28f-706e37191467 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:52.339589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:44.880939Z digest=sha256:f252fbdd398abd788ad88a86ab3f7f7d4e5df377c1c401a5309e9f174d662695

Observation c363396a-0a42-4644-b8a4-9fe8775af82d · outbound

This paper cites Chujie, S.

Accelerating RLHF Training with Reward Variance Increase Chujie, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.201484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:44.987046Z digest=sha256:4e98855738f426467d252d4f7303428ac154df98a4a999dde9658cc3f9fccab8

Observation b97eb7be-0ff2-497d-988d-50a8efd04ff0 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.988803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:45.089926Z digest=sha256:375e0d86ce73942950c5f2210d42fca253cf7531be09bfc6d100b38f92336bd6

Observation 87614c65-c5fd-4e18-bdb3-896c4ca1e193 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.807241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:45.227730Z digest=sha256:c15b33e0b609b721bdad536670b810e9eb01cd73d3a65e2986f70cc297c04781

Observation 997c44ba-346c-42fc-858a-3666c5699c72 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.534506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:45.334159Z digest=sha256:7f294dfdbf248e5b2fe35fe4fc997bc4b83ae866caae49c0f9046d21472444a5

Observation b1a7a1bb-a0c5-47c3-ae39-46bc7d4c07d0 · outbound

This paper cites Gr¨unbaum, V.

Accelerating RLHF Training with Reward Variance Increase Gr¨unbaum, V

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.371181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:45.466150Z digest=sha256:9c8a9fb9e577a55958d5638bdca88765dab0a31c2bf757220e5015c15a9b34a3

Observation e13f9c7f-146b-428b-ae70-b1175beb77b0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.553164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.553164Z digest=sha256:764d5c80100196f22c52c1b591f8783b522c6d6e74fa308382349cdcd1573baf

Observation 61f26049-22d0-4e45-9147-e926ac8b41d2 · outbound

This paper cites An Overview of Large Language Models for Statisticians.

Accelerating RLHF Training with Reward Variance Increase An Overview of Large Language Models for Statisticians

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.680041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.680041Z digest=sha256:0f896997e085a479384b7a062c395af4414c09752382e7d269de259ab4662789

Observation 38fc1af3-733a-47dc-80a4-1c3105557b5d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Accelerating RLHF Training with Reward Variance Increase RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.856856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.856856Z digest=sha256:f70e7ad047fe8bd1870738689cb60ae3865ea089ca19e1e57edbdc026061df51

Observation d7a1ad49-e044-4c75-b602-64ebd181ce73 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.973358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.973358Z digest=sha256:e93f484c3e6ba876754ebdacab2421eeb94e24b066a6758fb6c5c3e4e6002161

Observation 703955ed-8188-4880-bbc3-5b459386055e · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-V3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.078140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.078140Z digest=sha256:cbfc08bf953f658a2fe5216186d4dabcb73155d5dabec74cd7e0d2280055a09a

Observation 51262a20-3615-4009-b319-8df2320193ab · outbound

This paper cites How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?.

Accelerating RLHF Training with Reward Variance Increase How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.192236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.192236Z digest=sha256:c62bf3c4efee25eeeefb54029c1d13949dd71cae447f8906dd61d7a650a52913

Observation 2a18dfff-a744-402e-a18a-80706f1dd73e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Accelerating RLHF Training with Reward Variance Increase Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.332920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.332920Z digest=sha256:474d5a253148b99466dde90e86388e9f77c5d724db6a50fac712076954f39f23

Observation 3b06221c-1394-4095-9e18-8c0c2f95f4fd · outbound

This paper cites Ouyang, J.

Accelerating RLHF Training with Reward Variance Increase Ouyang, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:46.436502Z digest=sha256:95be9fc6c65bb238c56adbbf470752431ad1b8591e6bf5cc9288167e6d02319b

Observation 64e7dd7c-78dd-4bae-8327-e00e335e3915 · outbound

This paper cites Radford, K.

Accelerating RLHF Training with Reward Variance Increase Radford, K

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.956294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:46.506583Z digest=sha256:5070bcd65aff7826f5973becc4582213cc2f5ca24255637e3046657039b2d60a

Observation 75223c18-a2d3-4934-943e-de025d093d7c · outbound

This paper cites Rafailov, A.

Accelerating RLHF Training with Reward Variance Increase Rafailov, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.800422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:46.639735Z digest=sha256:e18fcca788c0c76f8fec16b2bf4b366a8d3500c14fec040de371002e1150f739

Observation cfbcce77-e685-4005-8254-0f9945216123 · outbound

This paper cites Razin, Z.

Accelerating RLHF Training with Reward Variance Increase Razin, Z

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.741897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.741897Z digest=sha256:ca3f1444731ceeacec2b67ccf714dd05e72d75d571989cc35f8188bc8cf707ec

Observation affcfb43-06b0-4c29-b9f0-a9397c72fec3 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Accelerating RLHF Training with Reward Variance Increase High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.842272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.842272Z digest=sha256:c670d0ab48e34efaf8796da8f4797b71c6e543723016cddf0e0ce8787f36c157

Observation f4287c15-8172-459b-863a-289302a6ce41 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Accelerating RLHF Training with Reward Variance Increase Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.991866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.991866Z digest=sha256:0b9e6a3bffe4f4e0a36328ef4e49ebd115d0c6a0598c4047d0b84cf62d7bafb2

Observation 119a85bf-f936-4372-865c-ad5fec2e324e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Accelerating RLHF Training with Reward Variance Increase DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.080154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.080154Z digest=sha256:3366246756b2d12a11768e49dc4397715b590c898e9e097c8f6f4f91d0165552

Observation ef203d19-aad7-492f-8405-b79b262047f3 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Accelerating RLHF Training with Reward Variance Increase Understanding the performance gap between online and offline alignment algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.229540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.229540Z digest=sha256:b6a409feaa706d12fc97196e6f46fdf0a6725be765a7d5f756dc0e01c005f224

Observation e7bb1405-a6c5-4f81-b012-9c8576dee5c9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Accelerating RLHF Training with Reward Variance Increase Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.359724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.359724Z digest=sha256:1c5ada7d7d32764994efadf253e3d70e3158a49af927f32b626480a22019b3f8

Observation 126dde1d-c5fd-4ec2-810a-70f7a6e82833 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating RLHF Training with Reward Variance Increase LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.506865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.506865Z digest=sha256:d1455c9328c95efa23f101c3ac3254c3fa700c4b7e05879a9b51c83dc251b09e

Observation c098633a-f25b-4ad0-bd18-8f66ba24e76e · outbound

This paper cites Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models.

Accelerating RLHF Training with Reward Variance Increase Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.666105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.666105Z digest=sha256:e71763fdc07698a1e4831ac46fde16cb82b70d4c40b1a0eb898e5e8ebb3f4ede

Observation 4be40304-6593-490b-9633-e738ab1320a2 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.775977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.775977Z digest=sha256:20205fdd1ad4f1868115d2c9d798d063d1b57aad55288a0b4048a24d5524d8fc

Observation ffdd6de7-4d7a-4546-80f4-97f8a1324b4b · outbound

This paper cites Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning.

Accelerating RLHF Training with Reward Variance Increase Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:49.583436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:47.876252Z digest=sha256:66fa64d005d6775d4da19c72e42fb1b2bc6e6174c5a1cf810e2a0794305f7379

Observation e5783d80-1ead-4e57-9135-51ec6639f0f9 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.617804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:48.013858Z digest=sha256:1fb1304f80037387c0d91b05e08f99a84d2fabdd0baf33106e3a7feda6e5a803

Observation 5a983e2c-f021-444c-964b-9088cb1550be · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.111769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.111769Z digest=sha256:22f551dd3952d0fd5a717afd3cd65d1acc404af9bed91f45cfbaf5bdef3a8a4d

Observation fb94bad4-608f-4d77-a0c3-982c36af98d4 · outbound

This paper cites Qwen3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.283646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.283646Z digest=sha256:7906e98434cab2b71b2b67c4bab6d8cf0ef872751c71474dde09b8280ea6c271

Observation 89d871bd-b7f4-4f56-9a54-030a5549edef · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.416612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:48.421561Z digest=sha256:e09d080f623642769a4656e47b8c66f1275ea0b73b8cd0423dec80bea2d818d2

Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.554997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.554997Z digest=sha256:5e58456551607ba506193882cffe730766816710021e6593465b048b5d0364ef

Observation 4bc96b5a-387e-48fa-9098-d040458b63c4 · outbound

This paper cites Zhang and C.

Accelerating RLHF Training with Reward Variance Increase Zhang and C

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.646323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.646323Z digest=sha256:64880666c4a3b0bef9a4deb839181759d3ff5c2ffc6954ac6279c47b8cdb6bfa

Observation 6882edae-8191-403f-89d4-5bfd3b936fea · outbound

This paper cites A Survey of Large Language Models.

Accelerating RLHF Training with Reward Variance Increase A Survey of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.798191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.798191Z digest=sha256:e848264192010d01bbe4386665436ef01420207f64315a21dd975be3d8c3a498

Observation b043d61d-bef4-4eca-8649-e5a056901647 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Accelerating RLHF Training with Reward Variance Increase Secrets of RLHF in Large Language Models Part I: PPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.911300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.911300Z digest=sha256:15ec19acc15e7f0615f4a60cbfb7e67d2668716ab8a4fa7152cf7dbbf1313de4

Observation 61a6078d-05bc-40bf-9c16-cd974259432c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Accelerating RLHF Training with Reward Variance Increase Fine-Tuning Language Models from Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:49.025901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:49.025901Z digest=sha256:9217ca6bd7d7be16e055b600e825d0084609e32989d3b79692386726ed4fef79

Observation 4d408927-a50c-40c7-8209-de3ac31fc322 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.328967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T12:59:49.168492Z digest=sha256:424693ac24028522316980e4f315133db93d33d4429035776b2ab290fa24ed91

Pith citing papers

No inbound Pith citation observations are available.