Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.927539Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2511.23310.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.927539Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0624fa21-0f17-4d6c-8d54-877a33b328a6 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc33d7b-b6a7-4c7d-ab3c-81ed1730e6ee · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf254ae-7ba7-41aa-a72d-801cfec47859 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cbba526-11e8-4349-b46e-84ff765fce1e · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a0004a-111e-42bf-928b-975c11a0dcb5 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c5c8dc-3789-4c5f-b578-07b39a20653d · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63dbecb6-46c8-41eb-bb11-94868fa0d084 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Deep reinforcement learning from human preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 661c9ad4-255c-49cb-9015-8a6eba4c3e1a · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e304497-bff5-4a3c-9e4b-d67f9b92166e · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Learning rate schedules for faster stochas- tic gradient search
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239f436c-ad32-4b16-93cc-578fe2bac4eb · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Note on learning rate schedules for stochastic optimiza- tion
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9a8a28-c780-4993-8e5e-d5e39d93e110 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b4841c-b26c-41d0-befa-183b775023e8 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Variance Reduction Tech- niques for Gradient Estimates in Reinforce- ment Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a002a902-d8b4-45b9-a9e1-a0cc5d4d337e · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4587cee8-b653-4aba-8d44-95657ecce8f5 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ∆LNormalization: Rethink Loss Aggregation in RL VR
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578bb1d3-8a0f-467f-8119-7cb6e6938399 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Measuring Mathematical Problem Solving With the MATH Dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1c747e-fd15-4ee6-bb25-3c02c9213ba0 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7b88388-7154-4c8e-a414-6fb6cdcf752d · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Buy 4 REINFORCE Samples, Get a Baseline for Free!
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f97b477-1bbf-47cc-bcec-5b7fc3c5ca07 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed123b43-a3e8-4bbe-9566-2c76e2078605 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards An Exponential Learning Rate Schedule for Deep Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f658452d-372c-446c-8f25-3536e859720a · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2481f592-cb45-4efc-8f9d-7efa6fdf1458 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Let's Verify Step by Step
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61bdd327-aa75-4922-91b9-d084dec88f66 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Hugging Face, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee1bef9-544c-4d3f-992d-ad87e0e32ad1 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Training language models to follow instructions with human feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca7a5e7-58a6-4d71-967f-36044d761b95 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd799f0-cf68-4120-830f-62c763e1bfe9 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c43904-127b-41c8-a4d9-a27bff661e3c · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d0d4d4-e43e-4743-b662-f81eecaeff01 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Proximal Policy Optimization Algorithms
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0a7930-a724-4ed5-a02c-e53dbd105095 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59837580-e41f-4789-bad2-050fb675b09e · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards HybridFlow: A Flexible and Efficient RLHF Framework
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6fa42e-8205-40e2-a885-07d7c53b4bff · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Solving Inequality Proofs with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d3676c-ed06-4713-afe5-3a48a78a2df4 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Learning to summarize from human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d354e495-9260-4fa8-a6ce-32bd312f9be2 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d833df9-5d94-440f-816d-05e97aca4710 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08af61d4-331e-4101-b547-749b90dc217c · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Self-Instruct: Aligning Language Models with Self-Generated Instruc- tions
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e46482-e302-406f-9ac5-dc8dc9079681 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346fc257-fbc5-4f41-afff-457ce96bdd5e · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Qwen3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0549afe-c503-48f6-b347-d34c7c8c723c · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d203758-1e76-44d3-9140-7655cad53b77 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e36d264-33fa-4622-8aa9-c856b9a9c289 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860f6358-57cb-4341-b488-26f773d3646a · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75231d1f-e143-44eb-8ebf-f0baf1334369 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a9f983f-0143-4fbe-99b7-51d84f4debc0 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20550033-098c-4543-b10d-e51c4554550c · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Formally, we consider min {Nt}T−1 t=0 ,{G t}T−1 t=0 E L(θ T ) s.t
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08aa9e3-008c-44d0-b1f5-b4efd877269b · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Practical recommendations for gradient-based training of deep architectures
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2ee59d-f3c4-464d-8a00-1d3a097ef619 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · outbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.