Pith. sign in

Paper Citation Record · LEDGER

Accelerating RLHF Training with Reward Variance Increase

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2505.23247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23247 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:49.168492Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77bbfb63-640e-41ac-880e-8354745ec0f1 · outbound

This paper cites GPT-4 Technical Report.

Accelerating RLHF Training with Reward Variance Increase GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.286507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.286507Z digest=sha256:9d73c834658fb2c1745146ece0929028e9c89a547244f3c1207242fa23cb2fbe

Observation 5cd1c3b0-8545-40f8-9e76-ab1e7b05cc95 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Accelerating RLHF Training with Reward Variance Increase Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.338823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.338823Z digest=sha256:666687d274ed4d9f53f3a46d201bd0d6eb020b0b9f93bb1685545b4fbd4680e0

Observation 5df03760-eb3a-4895-bc6e-e7617d7c849e · outbound

This paper cites Biderman, H.

Accelerating RLHF Training with Reward Variance Increase Biderman, H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.930185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:44.398065Z digest=sha256:6ec58a41237fc5a38ee5cdc417bf179601629123c95016da1c18c370fa7a4fa3

Observation 009e307f-77ea-4bc4-b906-2638e8a9c2b4 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Accelerating RLHF Training with Reward Variance Increase On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.493191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.493191Z digest=sha256:8a3744e0991cfb1797522386f1c8f89c1003a2b9e2848b4c02f34e1decafab78

Observation 873fd355-ecdc-4e96-bc55-42f8314b04fd · outbound

This paper cites Brown, B.

Accelerating RLHF Training with Reward Variance Increase Brown, B

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.702380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:44.558034Z digest=sha256:463dafd834e3bc5880fad96d03d4f857a4aebed1294a8f19bd3ff03c04c6438f

Observation 478a1110-21f8-48ec-80f5-02d5da0601ff · outbound

This paper cites Busa-Fekete, B.

Accelerating RLHF Training with Reward Variance Increase Busa-Fekete, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.492367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:44.669918Z digest=sha256:543a18062fda0e3353aed4a897d370ddc5bbcd982ec3a42d7888152779145a0e

Observation c28f5f59-6cf7-43b9-96d3-637382b39030 · outbound

This paper cites The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models.

Accelerating RLHF Training with Reward Variance Increase The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:50.084647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:44.776384Z digest=sha256:5fe0ecd4cc06a4a911d94be1b81178268c7e85949a50dff3b5b41a326e807b2e

Observation 54bf522b-4f9c-4680-a28f-706e37191467 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:52.339589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:44.880939Z digest=sha256:b79d90482eea551ac022e35176218be0f58ce14d13ff36e533d4dbe00925d8eb

Observation c363396a-0a42-4644-b8a4-9fe8775af82d · outbound

This paper cites Chujie, S.

Accelerating RLHF Training with Reward Variance Increase Chujie, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.201484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:44.987046Z digest=sha256:55f5cb641767febbddd82b88bfa1dab0ca55fa55fb48dddcdaff7f108a959ba8

Observation b97eb7be-0ff2-497d-988d-50a8efd04ff0 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.988803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:45.089926Z digest=sha256:db1368bfc0e1d36635436550899e85d2e35697fdb39f93e9e1d1b60a116987ff

Observation 87614c65-c5fd-4e18-bdb3-896c4ca1e193 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.807241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:45.227730Z digest=sha256:090286d85960c24d3fc5423e833f959b813d029eb993d96934a38f2af0138320

Observation 997c44ba-346c-42fc-858a-3666c5699c72 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.534506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:45.334159Z digest=sha256:9d9ef7a6a83b11412e269e8ce71614137c35ed17e796afb5cd17b13afd7d2485

Observation b1a7a1bb-a0c5-47c3-ae39-46bc7d4c07d0 · outbound

This paper cites Gr¨unbaum, V.

Accelerating RLHF Training with Reward Variance Increase Gr¨unbaum, V

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.371181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:45.466150Z digest=sha256:219089239de29de4530983ba045328bd56ad97c94e604923ca260617f243e0f3

Observation e13f9c7f-146b-428b-ae70-b1175beb77b0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.553164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.553164Z digest=sha256:68db8d4007f01e7d9cabc54d029bf67034d1601c4b64aac79a2c8d6fa71be8e3

Observation 61f26049-22d0-4e45-9147-e926ac8b41d2 · outbound

This paper cites An Overview of Large Language Models for Statisticians.

Accelerating RLHF Training with Reward Variance Increase An Overview of Large Language Models for Statisticians

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.680041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.680041Z digest=sha256:96a2efd6be43242be1c2facaa77fc5b07690bb1e89b08ae018345dddc617134e

Observation 38fc1af3-733a-47dc-80a4-1c3105557b5d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Accelerating RLHF Training with Reward Variance Increase RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.856856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.856856Z digest=sha256:c478a3b3fbb3f42e53e9cf13ca8c04e066c8e830c70338ee8a75b14441616088

Observation d7a1ad49-e044-4c75-b602-64ebd181ce73 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.973358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.973358Z digest=sha256:b1f651eef1a28e801fa6417540840a7a0c122fc12c7119a13f926f7f167d6378

Observation 703955ed-8188-4880-bbc3-5b459386055e · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-V3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.078140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.078140Z digest=sha256:840b2b7118a981c7f3ae4592da188d9bd5641f740b9a0be83aedbe2d8245858a

Observation 51262a20-3615-4009-b319-8df2320193ab · outbound

This paper cites How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?.

Accelerating RLHF Training with Reward Variance Increase How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.192236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.192236Z digest=sha256:7c6d0b3d7d2abec79fead8c531f79a4d8ae974d6f14717c79be35f742d19487c

Observation 2a18dfff-a744-402e-a18a-80706f1dd73e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Accelerating RLHF Training with Reward Variance Increase Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.332920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.332920Z digest=sha256:d774fe9bdfd4cd2831d486bee4f5b16e7fbeff4ab0e0bfdcd2073eb778691f26

Observation 3b06221c-1394-4095-9e18-8c0c2f95f4fd · outbound

This paper cites Ouyang, J.

Accelerating RLHF Training with Reward Variance Increase Ouyang, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:46.436502Z digest=sha256:5bfc49c558b402f2b41f8f96c461661367891a61f81c6138e52dbfc27b606be7

Observation 64e7dd7c-78dd-4bae-8327-e00e335e3915 · outbound

This paper cites Radford, K.

Accelerating RLHF Training with Reward Variance Increase Radford, K

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.956294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:46.506583Z digest=sha256:b144e152b5f945cf5e420746c6f3737bf929f30dfec58cb465abaa53e5a76457

Observation 75223c18-a2d3-4934-943e-de025d093d7c · outbound

This paper cites Rafailov, A.

Accelerating RLHF Training with Reward Variance Increase Rafailov, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.800422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:46.639735Z digest=sha256:c50ed51b0738bfd1fbfcef93547e4bd87a401d7d656dcfa3bc3c47dc7b0f0a52

Observation cfbcce77-e685-4005-8254-0f9945216123 · outbound

This paper cites Razin, Z.

Accelerating RLHF Training with Reward Variance Increase Razin, Z

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.741897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.741897Z digest=sha256:b8d4cc16ecf832f9f185c55554fb562ec41238473ef7f8d12819cea595b7ef7b

Observation affcfb43-06b0-4c29-b9f0-a9397c72fec3 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Accelerating RLHF Training with Reward Variance Increase High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.842272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.842272Z digest=sha256:b1d8d0be744ba80562c4c75b46881fb2a1d9624f5ff6776e7c4f680731a124ff

Observation f4287c15-8172-459b-863a-289302a6ce41 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Accelerating RLHF Training with Reward Variance Increase Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.991866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.991866Z digest=sha256:3217ea0bc333055f7c064aa8139156b1bc5d8ad105318cf489d07ce54473ad0c

Observation 119a85bf-f936-4372-865c-ad5fec2e324e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Accelerating RLHF Training with Reward Variance Increase DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.080154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.080154Z digest=sha256:a73e120a5d16bd0378ed7524f1e9fd47bbcd5f0cd0ea6936458f7d96a8ba8817

Observation ef203d19-aad7-492f-8405-b79b262047f3 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Accelerating RLHF Training with Reward Variance Increase Understanding the performance gap between online and offline alignment algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.229540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.229540Z digest=sha256:08ca25934bdaa4b61bbcc03746aa93ee0f2ff203904c3a765644e8b2ea395874

Observation e7bb1405-a6c5-4f81-b012-9c8576dee5c9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Accelerating RLHF Training with Reward Variance Increase Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.359724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.359724Z digest=sha256:8f3187edf8e24786d3726574bff9da0318cc97d694fea8286726c8b3653601d7

Observation 126dde1d-c5fd-4ec2-810a-70f7a6e82833 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating RLHF Training with Reward Variance Increase LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.506865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.506865Z digest=sha256:60e277b44ce2851a61dbed08efe253cda3ab1130433fc8defb8ed51c317d3360

Observation c098633a-f25b-4ad0-bd18-8f66ba24e76e · outbound

This paper cites Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models.

Accelerating RLHF Training with Reward Variance Increase Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.666105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.666105Z digest=sha256:77bc1d5d0850925a165b106ea32ec925b54e6354aa5952576f6411aefb331114

Observation 4be40304-6593-490b-9633-e738ab1320a2 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.775977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.775977Z digest=sha256:b7ead964d3298b5e5ad257458de6f024e0cc19d42e9a7c04cf92d6cfe86b460b

Observation ffdd6de7-4d7a-4546-80f4-97f8a1324b4b · outbound

This paper cites Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning.

Accelerating RLHF Training with Reward Variance Increase Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:49.583436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:47.876252Z digest=sha256:66c649e2ced503ae3748583598e904953f98a2758419549bc443c80d54f2ef47

Observation e5783d80-1ead-4e57-9135-51ec6639f0f9 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.617804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.013858Z digest=sha256:91f8feddaf3fdbb202c792db30d407f3f1cc577e8a568afd55078473419f4121

Observation 5a983e2c-f021-444c-964b-9088cb1550be · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.111769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.111769Z digest=sha256:183bfbe0d802013a17c7b4ecfa07dc173ca9f6296d862d5b5f623a5420ceb9d8

Observation fb94bad4-608f-4d77-a0c3-982c36af98d4 · outbound

This paper cites Qwen3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.283646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.283646Z digest=sha256:c60bb77e2af206fc546ec1b5050199baa8468a6a1f425ff30201260a5a6849a1

Observation 89d871bd-b7f4-4f56-9a54-030a5549edef · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.416612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.421561Z digest=sha256:4981dcb0ec17945f98c3b249cf0f703d0ca82b3f8327695598255fde95d842cb

Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.554997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.554997Z digest=sha256:57923003254ad683c88bda3b113925bd1090a6363e8219834de2581687ad87bf

Observation 4bc96b5a-387e-48fa-9098-d040458b63c4 · outbound

This paper cites Zhang and C.

Accelerating RLHF Training with Reward Variance Increase Zhang and C

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.646323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.646323Z digest=sha256:8702463fa62032125b4d0d89e71b90ebb13b51e43e05b2ed18c488920d2ccc25

Observation 6882edae-8191-403f-89d4-5bfd3b936fea · outbound

This paper cites A Survey of Large Language Models.

Accelerating RLHF Training with Reward Variance Increase A Survey of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.798191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.798191Z digest=sha256:1d35a950d19924a7baf2b1ad6b080411a435e4f28de9d6180ce28239e149753c

Observation b043d61d-bef4-4eca-8649-e5a056901647 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Accelerating RLHF Training with Reward Variance Increase Secrets of RLHF in Large Language Models Part I: PPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.911300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.911300Z digest=sha256:d26e2e16decfdf16c0955be3dff1084efe1cbd107e25946f1999db960858197d

Observation 61a6078d-05bc-40bf-9c16-cd974259432c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Accelerating RLHF Training with Reward Variance Increase Fine-Tuning Language Models from Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:49.025901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:49.025901Z digest=sha256:e86ba93acfe75e3ebba33ea70c078fe2c35f5bdf5bf2316509f93c110b1a2df5

Observation 4d408927-a50c-40c7-8209-de3ac31fc322 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.328967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.168492Z digest=sha256:1615cd928a7cd92335db1d64613659f6aba5dbbfd8bb8bce0eee9e0f0d5851e9

Pith citing papers

No inbound Pith citation observations are available.