Pith. sign in

Paper Citation Record · LEDGER

Accelerating RL for LLM Reasoning with Optimal Advantage Regression

As of 14 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 3 inbound Pith citation observations for arXiv:2505.20686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20686 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:57.728078Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.744697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:36:07.284835Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e81a2ab3-e7be-4a5c-b437-cf9bcd404833 · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:00.295544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.898779Z digest=sha256:1bed273b6181e3556c6b8ea36fec58e5ecac716ea7a43d44f97667708d6803e5

Observation 66aec46c-8ef6-45f2-8d3c-ebff8e96057e · outbound

This paper cites So,a2 = (−2)2 = 4.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,a2 = (−2)2 = 4

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.902129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.019407Z digest=sha256:21172e81bb21783fb56f8feb5d2b35fdd2fc5829d2532252812eab877424db35

Observation d84a8ba9-54ec-48b2-9408-3179fe2abd2e · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:58.688420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.481389Z digest=sha256:dfd0e7915e674b90fe47aee90ce24a949f6105559741ca9463417a2d4dc1a56a

Observation 325b6916-991d-49d0-b391-836cad9fdfb5 · outbound

This paper cites So,1 b2 = 1 32 = 1 9.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression So,1 b2 = 1 32 = 1 9

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:00.109577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:56.962513Z digest=sha256:3fc7688c1a0ebd87e69ed721778140fdd1a7a9ebd24c64c26c259e82ca648d0a

Observation 5dc3bda0-7723-4f97-9ff4-08425c39f29a · outbound

This paper cites Since3≤b≤5 , the minimum value ofb is 3.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since3≤b≤5 , the minimum value ofb is 3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.679835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.100364Z digest=sha256:37e371fc91b9aaa252a0a1b971b50acdc77a792e3031f720dd8f42af8821bdb9

Observation 4af27d99-0b7b-4316-86b3-b1c4283560c3 · outbound

This paper cites Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Since−6≤a≤ −2, the minimum value ofa2 is (−2)2 = 4

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:59.386030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.174578Z digest=sha256:62195a21f9855a188f3ab0f3e86747cbba942cba0ed6cec2aa3d1c716eead405

Observation 26adfdfb-51e0-40ad-97db-ff2df1d042ab · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:59.161666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.283784Z digest=sha256:d50ce3d5be0713938c3267a5980691ebe9e4babfeb0c505d820aa9c8300babb9

Observation 42880f05-e596-424c-9dd5-7c7c1c519283 · outbound

This paper cites an unresolved cited work.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:58.888017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.381757Z digest=sha256:cb6dd98b1df8fdaa1399df218ef3e11186e2b3b4a4be6bed71c90f02c5c6304d

Observation 8c2c594f-2fde-4bcc-9a59-d248968b7d93 · outbound

This paper cites From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression From these calculations, we see that the greatest possible value is indeed achieved whena=−2andb= 3, giving us: −35 9

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.506243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.546525Z digest=sha256:f00ce9d3b1f6b8b96d101a476a4de06935dbcd6c924aac0686de61f3a631e3a5

Observation f814d688-6ba3-4f94-b94b-90668a4e6144 · outbound

This paper cites a+c= 1−1 = 03.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression a+c= 1−1 = 03

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.263134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.653874Z digest=sha256:16734b4a7ddab2b5c70c5bb14d16e28c2b7c8bb0dd75f1c05b72973110516977

Observation 5ea03061-2779-4809-8854-e3be1289dffc · outbound

This paper cites X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression X y exp(⟨θT+1 , ϕ(x, y)⟩)P y′ exp(⟨θT+1 , ϕ(x, y′)⟩) − exp(⟨θ⋆, ϕ(x, y)⟩)P y′ exp(⟨θ⋆, ϕ(x, y′)⟩) # ≤Ex

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:58.018712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:55:57.728078Z digest=sha256:afd00e4db04fdd705bebb680521aa24ffc6d3ffd029b59db8ad91ab87bda61ff

Observation a25873b9-4bbb-48a7-b9d8-600f23a38b99 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.838631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.838631Z digest=sha256:21540aaf62827ad27d1b56042c0cb74caaa2880df473e28f0517ae5f65dbeef8

Observation 7031781b-ba7e-4080-88d5-327e559e5f6e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RL for LLM Reasoning with Optimal Advantage Regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.722789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.722789Z digest=sha256:e45bf6e40e331a47c6c431a2f6f7bdbe11a433ada54107072527ddaba0a5d44b

Pith citing papers

Observation 6349f898-48a5-4df7-a4d1-6431d7e0e02f · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.288018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:3b21675d3bebe304e788d03cfcef49749142bb864cbab4eff9045935cf048c45

Observation b6ff4adb-ef38-43aa-be2c-f7875be47739 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:19.340781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:19.340781Z digest=sha256:3137e4f3c469bfd951f460df7b885c7eaf9fd809614edcdd71d2ee5e401c7972

Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.744697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.744697Z digest=sha256:9049ace28cf159b0909eeaed91c40832df64220f01df2e1038f8a67a12d31e8a