Pith. sign in

Paper Citation Record · LEDGER

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2511.23310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.23310 v4

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.927539Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0624fa21-0f17-4d6c-8d54-877a33b328a6 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.715714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.715714Z digest=sha256:63dcfcdc54ccbf139e2872eed170cdc6c68f64c49d8bc89ca3e6e9405c7b019e

Observation cdc33d7b-b6a7-4c7d-ab3c-81ed1730e6ee · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.720985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.720985Z digest=sha256:c6c09e1458863e01af52ac654c0835b356e3338ba0e6f56b4a334326f815aea6

Observation adf254ae-7ba7-41aa-a72d-801cfec47859 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.725978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.725978Z digest=sha256:875aa4dfb2d834fc8f4025051c4a0456e64698900dfe05de84e08e6e69afc6c1

Observation 0cbba526-11e8-4349-b46e-84ff765fce1e · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.735535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.735535Z digest=sha256:d96e8f80f17a7a623109dfce8e923f375c55906d603e1b1f2ba9c3fa23545cf3

Observation 42a0004a-111e-42bf-928b-975c11a0dcb5 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.739630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.739630Z digest=sha256:5985cbdf20cd61505a0dc3fbd567a1794d5c8f84f6286c038248e0e00e5d9e4b

Observation b9c5c8dc-3789-4c5f-b578-07b39a20653d · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.749076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.749076Z digest=sha256:9da4cd8be177fb04a69082974c36097f6fbf04c1fe909b943a399f5a68f75008

Observation 63dbecb6-46c8-41eb-bb11-94868fa0d084 · outbound

This paper cites Deep reinforcement learning from human preferences.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.753929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.753929Z digest=sha256:eadb99ad993316bd67eccc96e998a1d55902a8312ade7cd65ba77dc4ded98510

Observation 661c9ad4-255c-49cb-9015-8a6eba4c3e1a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.758362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.758362Z digest=sha256:d456ec80338ff64e6b74ee23a5bef578f0781d1e9a3cc62067da383358ca4e20

Observation 0e304497-bff5-4a3c-9e4b-d67f9b92166e · outbound

This paper cites Learning rate schedules for faster stochas- tic gradient search.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Learning rate schedules for faster stochas- tic gradient search

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.763412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.763412Z digest=sha256:8e7140deb00c8456107ccdcd2be49c0f452d88a4f53de217c0dac4f918031497

Observation 239f436c-ad32-4b16-93cc-578fe2bac4eb · outbound

This paper cites Note on learning rate schedules for stochastic optimiza- tion.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Note on learning rate schedules for stochastic optimiza- tion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.768142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.768142Z digest=sha256:2a2e8ab58870a61d763b6a4116017319b3d343455fb08d0fd1540a2fa20752c7

Observation 1f9a8a28-c780-4993-8e5e-d5e39d93e110 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.772206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.772206Z digest=sha256:b636bfe0720001e232d8c13cfbdcdb49a89377e1c7a01a5e7c88ce560d83f7f0

Observation c0b4841c-b26c-41d0-befa-183b775023e8 · outbound

This paper cites Variance Reduction Tech- niques for Gradient Estimates in Reinforce- ment Learning.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Variance Reduction Tech- niques for Gradient Estimates in Reinforce- ment Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.776040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.776040Z digest=sha256:cee90c8e371b09a2ea5484a3418a83bbd138a7b469560f6029302c51013eb8a2

Observation a002a902-d8b4-45b9-a9e1-a0cc5d4d337e · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.781021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.781021Z digest=sha256:0581f8711b3f33c75609ebd90214325091c367cd0394378bd7418147a11f585c

Observation 4587cee8-b653-4aba-8d44-95657ecce8f5 · outbound

This paper cites ∆LNormalization: Rethink Loss Aggregation in RL VR.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ∆LNormalization: Rethink Loss Aggregation in RL VR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.785322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.785322Z digest=sha256:b2ecfeb37d50d25e0a9a4931d80a60992d1c9db1d133731c8df5b402dbbf69cd

Observation 578bb1d3-8a0f-467f-8119-7cb6e6938399 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.790052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.790052Z digest=sha256:8229d7c13369ddaaa938a1157d5315c170ac5dc36e59e6e29786bcb749a3e257

Observation 0e1c747e-fd15-4ee6-bb25-3c02c9213ba0 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.794213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.794213Z digest=sha256:c9ab6aabd12799c236a434f5b417cc4f1c22887c38c9ed79f2126d1e155dae1d

Observation d7b88388-7154-4c8e-a414-6fb6cdcf752d · outbound

This paper cites Buy 4 REINFORCE Samples, Get a Baseline for Free!.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Buy 4 REINFORCE Samples, Get a Baseline for Free!

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.798542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.798542Z digest=sha256:4439e322fc18c9576472c283ca3ab1f720b5a240d78d75d9402f9bae794d7803

Observation 6f97b477-1bbf-47cc-bcec-5b7fc3c5ca07 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.802389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.802389Z digest=sha256:75191f558c1facc090a9cf7646dfa56f94394c6e02d2e916b66e4254c301ef1e

Observation ed123b43-a3e8-4bbe-9566-2c76e2078605 · outbound

This paper cites An Exponential Learning Rate Schedule for Deep Learning.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards An Exponential Learning Rate Schedule for Deep Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.806756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.806756Z digest=sha256:5fd118fe8c707abc03113945a7fd29265232592740b81a38280d9236ed76533b

Observation f658452d-372c-446c-8f25-3536e859720a · outbound

This paper cites Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.811511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.811511Z digest=sha256:be51cae19b5ed3c20745cad866aec325474f5a0f9e51f59ba0beb90f19fb9b6a

Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.815569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.815569Z digest=sha256:ef92f456cfda9ccf1a1bf2748497e8e08ef72c7fbc184c3afb53f9590bb20969

Observation 2481f592-cb45-4efc-8f9d-7efa6fdf1458 · outbound

This paper cites Let's Verify Step by Step.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Let's Verify Step by Step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.820370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.820370Z digest=sha256:225603df8c127e96bf1f440484c33252ee6f2306a4baffcb435433f6e8fe260a

Observation 61bdd327-aa75-4922-91b9-d084dec88f66 · outbound

This paper cites Hugging Face, 2023.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Hugging Face, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.824626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.824626Z digest=sha256:f7b34c044940948bc7e044b24fa080dbabb6b671a4aed5b49b04cef77cb923db

Observation 5ee1bef9-544c-4d3f-992d-ad87e0e32ad1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Training language models to follow instructions with human feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.829212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.829212Z digest=sha256:bfc1d245cca7e5f1bb82adebe4b146229bf5c16fe7e2628284d6c1c605fb62af

Observation 6ca7a5e7-58a6-4d71-967f-36044d761b95 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.833625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.833625Z digest=sha256:7757863a4c41eb19458d87bd50e7b5ed6c0e775131d4cc16cf211220850c40e9

Observation fdd799f0-cf68-4120-830f-62c763e1bfe9 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.838194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.838194Z digest=sha256:9ddd9d6ed516c11acb8b0df4376de6f211ca6e830a75c8833b8a0ab459436981

Observation 53c43904-127b-41c8-a4d9-a27bff661e3c · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.842538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.842538Z digest=sha256:34487a350cb490bb3f950daf2c34399037f4848ed1a023cc68997ddd78c955ab

Observation b2d0d4d4-e43e-4743-b662-f81eecaeff01 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.851666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.851666Z digest=sha256:107ff51fbb713d8db5c325395286673a7a29dcbf1740527163c2423d9e46fbb2

Observation ce0a7930-a724-4ed5-a02c-e53dbd105095 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.856427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.856427Z digest=sha256:7b427c5614296804cdc800fafff29a4b5767fcdf167e413773e25f6f89121410

Observation 59837580-e41f-4789-bad2-050fb675b09e · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards HybridFlow: A Flexible and Efficient RLHF Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.860531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.860531Z digest=sha256:c01719e3122b894f696ee8c34feb2e72af23b302f87852ee5c0af20e94f7bc55

Observation 6f6fa42e-8205-40e2-a885-07d7c53b4bff · outbound

This paper cites Solving Inequality Proofs with Large Language Models.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Solving Inequality Proofs with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.865402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.865402Z digest=sha256:279abb1d3129113fe4aeab40f2dc6d5423793899eb6cda03c4159aa8eae2f255

Observation 03d3676c-ed06-4713-afe5-3a48a78a2df4 · outbound

This paper cites Learning to summarize from human feedback.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Learning to summarize from human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.869562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.869562Z digest=sha256:db385d9cc4a1e2e675ab448e02736e3abb9ac0f79918bd6be756f0b242f3dcbf

Observation d354e495-9260-4fa8-a6ce-32bd312f9be2 · outbound

This paper cites Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.873682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.873682Z digest=sha256:327f5d36c4cca2cb0f967e52f5365a71dc78e9bdadf6b4421c9cb8fe7903ef7b

Observation 8d833df9-5d94-440f-816d-05e97aca4710 · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.878594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.878594Z digest=sha256:98e2ef0f1fd7cbf50665606987e181d93339e646637671f6fd73b38631a69c7f

Observation 08af61d4-331e-4101-b547-749b90dc217c · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instruc- tions.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Self-Instruct: Aligning Language Models with Self-Generated Instruc- tions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.883512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.883512Z digest=sha256:d270d1a7179dde3db959a4688634878dd17c3144da85ac0a16760d42889f6e6d

Observation a8e46482-e302-406f-9ac5-dc8dc9079681 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.887324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.887324Z digest=sha256:23815cd8631a41a2c1e18057bdcaf0f695020b736be9509f9a2c277d98018bed

Observation 346fc257-fbc5-4f41-afff-457ce96bdd5e · outbound

This paper cites Qwen3 Technical Report.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.896488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.896488Z digest=sha256:63014b1216dfda12fb5f2184438dae33d10f9b968260b0db56ea029686b54dcf

Observation f0549afe-c503-48f6-b347-d34c7c8c723c · outbound

This paper cites Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.900557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.900557Z digest=sha256:9a2e4c6c32159896f5186cd05a5f302bb7943339c3644afd0a3328f18dc3eed7

Observation 5d203758-1e76-44d3-9140-7655cad53b77 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.891598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.891598Z digest=sha256:2eccf2ecec52ecedc5b329fd1ae1b10431cde8d787e05d626fc2d8f53a96d098

Observation 7e36d264-33fa-4622-8aa9-c856b9a9c289 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.909663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.909663Z digest=sha256:164a5a1ff08b679f36ad3910140ef39a829caad770ed0aa51c2df4b58f88cdec

Observation 860f6358-57cb-4341-b488-26f773d3646a · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.913719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.913719Z digest=sha256:aac68d9314b63c82fd8b31f718410d636e60ec30c4d8e849ad32cfaa99bf9598

Observation 75231d1f-e143-44eb-8ebf-f0baf1334369 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.905721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.905721Z digest=sha256:017cb349859668d48fdfa06fc1f32e4fa775c073d7d7fe3fc24d0a1c91668ca5

Observation 9a9f983f-0143-4fbe-99b7-51d84f4debc0 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.918458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.918458Z digest=sha256:945d7d88d4f0ac11f60e5f43b6c69300284acf377450c2ec9afd0acf88d63945

Observation 20550033-098c-4543-b10d-e51c4554550c · outbound

This paper cites Formally, we consider min {Nt}T−1 t=0 ,{G t}T−1 t=0 E L(θ T ) s.t.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Formally, we consider min {Nt}T−1 t=0 ,{G t}T−1 t=0 E L(θ T ) s.t

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.927539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.927539Z digest=sha256:1326d4f9875ded2e462192aee6c82144fae602c3e70fe6e8a4dce9c0519f152d

Observation e08aa9e3-008c-44d0-b1f5-b4efd877269b · outbound

This paper cites Practical recommendations for gradient-based training of deep architectures.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Practical recommendations for gradient-based training of deep architectures

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.730490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.730490Z digest=sha256:7d712da446db0d32167ca4b8c77b2ab9236485213d61cc5c7ee522367a876f00

Observation 0d2ee59d-f3c4-464d-8a00-1d3a097ef619 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.922717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.922717Z digest=sha256:dd915c1d3df5a2c2c7feeaf262c774d9808a1632c3d6ed98ddf0b3393a894f34

Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · outbound

This paper cites Accelerating RL for LLM Reasoning with Optimal Advantage Regression.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.744697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.744697Z digest=sha256:576eebc851488f5d845ae352aad6240936f08c6e0a4c879cd8599b09eb5806ad

Pith citing papers

No inbound Pith citation observations are available.