Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:46.727588Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 5 inbound Pith citation observations for arXiv:2508.13023.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:46.727588Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:54:32.886708Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T19:45:00.969348Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c75a34cc-36a6-4bca-8bd6-59520750e758 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Online difficulty filtering for reasoning oriented reinforcement learning, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 149cde66-37c7-40ba-8ba9-98ba139589a0 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Evaluating Large Language Models Trained on Code
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1b8c73-9d9c-4ffa-84f1-a3be7ea22c59 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance SFT memorizes, RL generalizes: A comparative study of foundation model post-training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45835a5f-e947-45ab-af7b-d69fe5bd0fdc · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Process reinforcement through implicit rewards
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75ccf331-5611-4755-91d6-4b8a524c23c2 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Reinforcement learning for reasoning in small llms: What works and what doesn't, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e188b9d9-5485-4f0e-92a6-4625d284af3b · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2bc58f9-b8bc-4227-bb8f-931ed576b5ea · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance rstar-math: Small LLM s can master math reasoning with self-evolved deep thinking
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7a1b0738-3d2c-4cfb-a326-994cb350657b · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Deepseek-coder: When the large language model meets programming-the rise of code intelligence
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a0613843-62b1-4801-8b74-e0ea1ba4c4d2 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e79b2e2-ae76-46b7-972d-c330b1460853 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Measuring mathematical problem solving with the MATH dataset
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 71146fb6-1d70-4372-b783-973735f60fef · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Efficoder: Enhancing code generation in large language models through efficiency-aware fine-tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 515f637c-d318-4d4b-881e-95ea3ddf5c92 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6520eb37-07a4-462b-8a0a-59f6bc70fdab · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Openai o1 system card
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f048cd56-dd51-4081-8bd0-cb5ab0482af2 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c6b147c-abe7-4112-98d1-58c6628e0006 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde8c8b0-9c50-4718-8f10-b7d7b5771721 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Large language models are zero-shot reasoners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda7d050-c78a-431a-a6d6-2fd821607e3a · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Token-supervised value models for enhancing mathematical reasoning capabilities of large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d7ac7730-610d-40cf-bb66-88d44aedb63d · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Solving quantitative reasoning problems with language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879fe752-d194-424d-92c7-328228aa2cfb · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Adaptive group policy optimization: Towards stable training and token-efficient reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c516c940-adde-42ac-8535-18ab9e8faf8e · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 243c78a3-5ade-4f3d-b10a-0bff989019c9 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance ToRL: Scaling Tool-Integrated RL
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83670a95-f09e-4169-9018-1c0fe7410c3b · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Cppo: Accelerating the training of group relative policy optimization-based reasoning models, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd9b5000-af6d-42ab-8f6f-1b066d270705 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Understanding R1-Zero-Like Training: A Critical Perspective
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827b1cfb-9337-42af-b588-a928234e0c02 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Small language models: Survey, measurements, and insights
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e7fe4ed3-ea77-4cd6-a2b2-09589415a285 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance s1: Simple test-time scaling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1b26f9eb-f1dd-44bc-b93e-f58c7b7dbc54 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance s1: Simple test-time scaling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daff95ec-0fc9-4ba0-9be5-c2f08db63c37 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7534cf-b858-42f5-90f0-b47a108a5dc3 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A Survey of Small Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96295283-9d95-404b-b7c7-15f5b92c4957 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Curriculum reinforcement learning from easy to hard tasks improves llm reasoning, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3cd0096-4b07-4adb-9b1f-a4335e08728d · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d861c179-3a45-4ee4-b0c5-ebe061af31bc · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Gpqa: A graduate-level google-proof q&a benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43617af5-1158-4521-af8e-171a75d3782e · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Proximal Policy Optimization Algorithms
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f984c72-a38c-4164-b1ed-eeee9ba464cd · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdfee086-c493-47ad-a94c-e0a00f804dfe · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953ad0cb-4862-4553-a3fd-c6e493995ea5 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Code generation with small language models: A deep evaluation on codeforces, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4733867b-9ed2-46e3-a54b-40009ccd5379 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Qwq: Reflect deeply on the boundaries of the unknown
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76e8ca0-1475-4b0d-8565-7ec62200a0e2 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ea1174-4c35-4ef0-931f-4ffa6b18f3df · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Chain-of-thought prompting elicits reasoning in large language models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b49273bd-e940-414d-b5ec-bce514bfdff2 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b68e03fb-8021-4f54-b2a5-cb5efbcfa408 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea8d5ff3-2a6b-4bcd-8af6-f078c415f734 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Rlvr-world: Training world models with reinforcement learning, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cb60d8-9c69-45c1-8da3-8c4b6d8d5319 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2536eb08-d3cc-4ee0-a53d-7c14dec8db69 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a364484-5e2a-491d-9356-873fcff38a0d · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 349fc563-0175-4f33-b289-f000cc71b7a3 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Qwen3 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a147ee5-e9af-423c-bcb8-e23668805dce · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Treerpo: Tree relative policy optimization, 2025 b
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca10bc2-ba30-4015-879a-2d47cc5917f3 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance LIMO: Less is More for Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6909db26-c4a7-4eba-b806-47763d3c1d22 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Dapo: An open-source llm reinforcement learning system at scale
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7e23046-74c3-432f-bca3-3e320e6856b8 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77fd2a2b-5bef-42fc-aafd-5b6f5da33df2 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d084dc3c-910a-4cbc-9535-8d8786694e82 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc47b08-9b7a-4a03-80dd-de0235f3c273 · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Disco balances the scales: Adaptive domain- and difficulty-aware reinforcement learning on imbalanced data, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a779c09c-06eb-4a63-97fe-ad541edafe3d · outbound
G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A technical study into 0.5b reasoning language models, 2025
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c77f7630-e7d6-44fb-b63a-597701da933d · inbound
A Survey of Reinforcement Learning for Large Reasoning Models G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7199d10d-8201-4edd-bc17-a5419835fcc6 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 965b4760-9fc6-4212-9403-9f045c9e2cf2 · inbound
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 352b23a7-b5db-4e71-9ddd-108c87290dcf · inbound
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a7c748f7-a37a-483c-9dcc-c030f20fa417 · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.