Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.728078Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 3 inbound Pith citation observations for arXiv:2505.20686.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.728078Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.744697Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T08:36:07.284835Z
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e81a2ab3-e7be-4a5c-b437-cf9bcd404833 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66aec46c-8ef6-45f2-8d3c-ebff8e96057e · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,a2 = (−2)2 = 4
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d84a8ba9-54ec-48b2-9408-3179fe2abd2e · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 325b6916-991d-49d0-b391-836cad9fdfb5 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,1 b2 = 1 32 = 1 9
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5dc3bda0-7723-4f97-9ff4-08425c39f29a · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since3≤b≤5 , the minimum value ofb is 3
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4af27d99-0b7b-4316-86b3-b1c4283560c3 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26adfdfb-51e0-40ad-97db-ff2df1d042ab · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 42880f05-e596-424c-9dd5-7c7c1c519283 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c2c594f-2fde-4bcc-9a59-d248968b7d93 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f814d688-6ba3-4f94-b94b-90668a4e6144 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression a+c= 1−1 = 03
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ea03061-2779-4809-8854-e3be1289dffc · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a25873b9-4bbb-48a7-b9d8-600f23a38b99 · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7031781b-ba7e-4080-88d5-327e559e5f6e · outbound
Accelerating RL for LLM Reasoning with Optimal Advantage Regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6349f898-48a5-4df7-a4d1-6431d7e0e02f · inbound
On the optimization dynamics of RLVR: Gradient gap and step size thresholds Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6ff4adb-ef38-43aa-be2c-f7875be47739 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.