Pith. sign in

Paper Citation Record · LEDGER

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

As of 17 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 5 inbound Pith citation observations for arXiv:2508.13023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13023 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:46.727588Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:54:32.886708Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:45:00.969348Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c75a34cc-36a6-4bca-8bd6-59520750e758 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Online difficulty filtering for reasoning oriented reinforcement learning, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.505625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.505625Z digest=sha256:dee82412c6a766b5658995c3fe04cf97cd0d54f8cf252e6941f23f59e1fdee27

Observation 149cde66-37c7-40ba-8ba9-98ba139589a0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.511599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.511599Z digest=sha256:c9b6666507e580b550cdfa5872a3bd68b0ef7cfa2adcd74721c705e9388a4e88

Observation 8a1b8c73-9d9c-4ffa-84f1-a3be7ea22c59 · outbound

This paper cites SFT memorizes, RL generalizes: A comparative study of foundation model post-training.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance SFT memorizes, RL generalizes: A comparative study of foundation model post-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.144803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.518887Z digest=sha256:cb6352702d2364522ba206e05e98a9796456d117a027f8e612f0bb1c5340f528

Observation 45835a5f-e947-45ab-af7b-d69fe5bd0fdc · outbound

This paper cites Process reinforcement through implicit rewards.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Process reinforcement through implicit rewards

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.131287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.524548Z digest=sha256:2540a57d14282fbe5a3d12d13764a2abe141228f32aceb9eaa20cac4bd5725b4

Observation 75ccf331-5611-4755-91d6-4b8a524c23c2 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn't, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Reinforcement learning for reasoning in small llms: What works and what doesn't, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.529591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.529591Z digest=sha256:53c45add16773e79dfb2c40ad054a649afc6d50929424e329bf3402f1ce02ac9

Observation e188b9d9-5485-4f0e-92a6-4625d284af3b · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.535537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.535537Z digest=sha256:e09b7b111139df3e97ca245cf5aa6cef21d29cd053fa42d691f21f1f112e08ec

Observation c2bc58f9-b8bc-4227-bb8f-931ed576b5ea · outbound

This paper cites rstar-math: Small LLM s can master math reasoning with self-evolved deep thinking.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance rstar-math: Small LLM s can master math reasoning with self-evolved deep thinking

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.106451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.540586Z digest=sha256:273de5bafe2e97a45f889004c627c608ff8f2fd062355560245d8ec7a4c87af8

Observation 7a1b0738-3d2c-4cfb-a326-994cb350657b · outbound

This paper cites Deepseek-coder: When the large language model meets programming-the rise of code intelligence.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Deepseek-coder: When the large language model meets programming-the rise of code intelligence

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.091920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.545294Z digest=sha256:5a0b049e66ca03a5d301ee993e74ef3d8f07559b54772fc79a41831b9c970b5d

Observation a0613843-62b1-4801-8b74-e0ea1ba4c4d2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.549537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.549537Z digest=sha256:92c2e12944a77972fd9a5fee1ecc19c0ae430dd86d2b76ce638eecea4322fbac

Observation 6e79b2e2-ae76-46b7-972d-c330b1460853 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Measuring mathematical problem solving with the MATH dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.079391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.553777Z digest=sha256:830c71f2db14cf0c0eceb0f54c8b9d5157f1b8446973d081ecf7409760643f86

Observation 71146fb6-1d70-4372-b783-973735f60fef · outbound

This paper cites Efficoder: Enhancing code generation in large language models through efficiency-aware fine-tuning.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Efficoder: Enhancing code generation in large language models through efficiency-aware fine-tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.064375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.558402Z digest=sha256:96eac207a52687230b5fed7ce82d0b2f379ca684070a4b0dd392928afad642af

Observation 515f637c-d318-4d4b-881e-95ea3ddf5c92 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.562236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.562236Z digest=sha256:7660d23735dce478dbe7e65697b786d338196c990ddc0b92f1e536e98f72ce90

Observation 6520eb37-07a4-462b-8a0a-59f6bc70fdab · outbound

This paper cites Openai o1 system card.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Openai o1 system card

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.048924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.566203Z digest=sha256:df3077afe3f818ebf2955347d4678bc629f75d2919a948275495a2f4909a2381

Observation f048cd56-dd51-4081-8bd0-cb5ab0482af2 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.034588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.570209Z digest=sha256:596589f7f5fa4e52160bf25a6e6e422d6c008e3475c57f3d25936855c5ee927b

Observation 0c6b147c-abe7-4112-98d1-58c6628e0006 · outbound

This paper cites Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.573914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.573914Z digest=sha256:24296b3580aab99dff02ef4c8ce93c7e4807ec82734e47e1c48a0784d0594520

Observation fde8c8b0-9c50-4718-8f10-b7d7b5771721 · outbound

This paper cites Large language models are zero-shot reasoners.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Large language models are zero-shot reasoners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.577908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.577908Z digest=sha256:9502fba3b2a40ebd7fc915d6c19ab2f7c8f195bb48603275838f5adca5492e69

Observation bda7d050-c78a-431a-a6d6-2fd821607e3a · outbound

This paper cites Token-supervised value models for enhancing mathematical reasoning capabilities of large language models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Token-supervised value models for enhancing mathematical reasoning capabilities of large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:48.014470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.581529Z digest=sha256:7b4649ebbfdafaa22af73f6a13c1b9a2a482f315cb54f808d2dd702a7cd9334d

Observation d7ac7730-610d-40cf-bb66-88d44aedb63d · outbound

This paper cites Solving quantitative reasoning problems with language models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Solving quantitative reasoning problems with language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.585647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.585647Z digest=sha256:44d959cd15da25947085e83631334b09d523cb8f772157541cbb56ae5f78c97c

Observation 879fe752-d194-424d-92c7-328228aa2cfb · outbound

This paper cites Adaptive group policy optimization: Towards stable training and token-efficient reasoning.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Adaptive group policy optimization: Towards stable training and token-efficient reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.589861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.589861Z digest=sha256:b7c4ca93a7e892690c161b1c6e9731cf7bbec42d49e827684d8f50618db6407b

Observation c516c940-adde-42ac-8535-18ab9e8faf8e · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions, 2024.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:47.993545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.593437Z digest=sha256:17725bc6b0f2e052d0fe89e8c4a61e1b663d6214873ba4bafef9ca497577995e

Observation 243c78a3-5ade-4f3d-b10a-0bff989019c9 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance ToRL: Scaling Tool-Integrated RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.597081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.597081Z digest=sha256:8da79ad54d112457c440dc6d1d81a164847f862b6c3c04c6d4402ea58310228d

Observation 83670a95-f09e-4169-9018-1c0fe7410c3b · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Cppo: Accelerating the training of group relative policy optimization-based reasoning models, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.602068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.602068Z digest=sha256:339da90c74b104bce5207a1c609c7e102a5b77bce540a1b53510ddb8fd9acd10

Observation cd9b5000-af6d-42ab-8f6f-1b066d270705 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Understanding R1-Zero-Like Training: A Critical Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.608105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.608105Z digest=sha256:c307e008209e7e15b3cbc42aa91d2beac9d786ba97ba4eea40059ca664ef9e6f

Observation 827b1cfb-9337-42af-b588-a928234e0c02 · outbound

This paper cites Small language models: Survey, measurements, and insights.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Small language models: Survey, measurements, and insights

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:47.980090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.612906Z digest=sha256:1f6647492959b737a1640a1883a27d710d1dbddf987ba3b4e76a01e967031e44

Observation e7fe4ed3-ea77-4cd6-a2b2-09589415a285 · outbound

This paper cites s1: Simple test-time scaling.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance s1: Simple test-time scaling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:47.967672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.616508Z digest=sha256:c1c78078477d864d7ada23c90694ba7eee2cae6604ad4691c70b362bef8a700d

Observation 1b26f9eb-f1dd-44bc-b93e-f58c7b7dbc54 · outbound

This paper cites s1: Simple test-time scaling.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance s1: Simple test-time scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.620199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.620199Z digest=sha256:672e6f6c8e3813a940655f60180d8501a25697596ff862a90235b89f856acd2f

Observation daff95ec-0fc9-4ba0-9be5-c2f08db63c37 · outbound

This paper cites Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.624230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.624230Z digest=sha256:b7c1504ad91513c01c35caa1db8a0e1978623d2bbf1e3f4fad7698f070047db8

Observation cb7534cf-b858-42f5-90f0-b47a108a5dc3 · outbound

This paper cites A Survey of Small Language Models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A Survey of Small Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.628011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.628011Z digest=sha256:0c9cba51a4b78f169df9f934a36cb57609d33542f8c1e252504f34001741f8e2

Observation 96295283-9d95-404b-b7c7-15f5b92c4957 · outbound

This paper cites Curriculum reinforcement learning from easy to hard tasks improves llm reasoning, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Curriculum reinforcement learning from easy to hard tasks improves llm reasoning, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.632022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.632022Z digest=sha256:78544fc190cce3ff0d5836c068cfc8927ca6e3873fa4d2b9d0bff8db94a3e84e

Observation e3cd0096-4b07-4adb-9b1f-a4335e08728d · outbound

This paper cites an unresolved cited work.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.635814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.635814Z digest=sha256:cff280f47f9616ad3776f4fc39926c3253ad831923169984b8b3e93fc506d0c7

Observation d861c179-3a45-4ee4-b0c5-ebe061af31bc · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Gpqa: A graduate-level google-proof q&a benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.640606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.640606Z digest=sha256:ef9b59d6e9bc008e47f856132b222a67a54fc7e0ef49c61500aad9301057b1d8

Observation 43617af5-1158-4521-af8e-171a75d3782e · outbound

This paper cites Proximal Policy Optimization Algorithms.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.644451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.644451Z digest=sha256:a72c6d45771f6d2600d6d613f62e6bb8b7baf6c41d7ab8232f93777e22baadef

Observation 8f984c72-a38c-4164-b1ed-eeee9ba464cd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.648332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.648332Z digest=sha256:2a37aabc9a3cc614389e658fcfa310dee6dd2ee033d6a085747f18231092c665

Observation fdfee086-c493-47ad-a94c-e0a00f804dfe · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.652856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.652856Z digest=sha256:9e3cabbef0a77167f9f6c2fcec7b5525cfd8a63c1c3c65e8dc7f3e54fa904212

Observation 953ad0cb-4862-4553-a3fd-c6e493995ea5 · outbound

This paper cites Code generation with small language models: A deep evaluation on codeforces, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Code generation with small language models: A deep evaluation on codeforces, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.656934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.656934Z digest=sha256:326c04083187e66b793e2bb6a148df93324215618d79bf46c5015663871acbe6

Observation 4733867b-9ed2-46e3-a54b-40009ccd5379 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Qwq: Reflect deeply on the boundaries of the unknown

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.660663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.660663Z digest=sha256:c404fbeffdd0f69fdcf42b33bff4380d9555d56a9b9e21ff08b56435b9292101

Observation b76e8ca0-1475-4b0d-8565-7ec62200a0e2 · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.664569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.664569Z digest=sha256:ef47a31e94aaecc6e29ebc9f9a41ffcc2b5f2a68c699476ec1f8cfefda3986a0

Observation d6ea1174-4c35-4ef0-931f-4ffa6b18f3df · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Chain-of-thought prompting elicits reasoning in large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.668507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.668507Z digest=sha256:444ea66a6b9f8ff68f3abb1ae5aefe579eecd74b82472d80fee9950246689fc4

Observation b49273bd-e940-414d-b5ec-bce514bfdff2 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.672374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.672374Z digest=sha256:2153bb01d812fe866dfbf2d16a98e33b5d99ae19f7aa2d1924a01550b5fff91b

Observation b68e03fb-8021-4f54-b2a5-cb5efbcfa408 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.676046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.676046Z digest=sha256:8786700fd01eac332d65296c6708a8083316f5cec34d825f230ea0ac3f7f5cdf

Observation ea8d5ff3-2a6b-4bcd-8af6-f078c415f734 · outbound

This paper cites Rlvr-world: Training world models with reinforcement learning, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Rlvr-world: Training world models with reinforcement learning, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.680719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.680719Z digest=sha256:33efc782f2e300acebfa25d3e7e70d5bdea000f4813366b6270047c31e2ddf93

Observation 30cb60d8-9c69-45c1-8da3-8c4b6d8d5319 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.684636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.684636Z digest=sha256:9f2d3222ac982a57768f500dbe85da96f4ef939dfe130c83380907c167ad60ac

Observation 2536eb08-d3cc-4ee0-a53d-7c14dec8db69 · outbound

This paper cites Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.688754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.688754Z digest=sha256:f8fb417ffecc467d6fd3dc5fd2512dad87eb9909fa6cec25ab5c7e9508c4813d

Observation 8a364484-5e2a-491d-9356-873fcff38a0d · outbound

This paper cites an unresolved cited work.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:19:47.930327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.692559Z digest=sha256:8e1ad87e4b63175e6abd4e7359b704f49937d1452a19f2515ae292a3ba60c06e

Observation 349fc563-0175-4f33-b289-f000cc71b7a3 · outbound

This paper cites Qwen3 Technical Report.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.696165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.696165Z digest=sha256:733e17b9ac410965f1e0309a325a5af1b8ad31cf617e85d09d86e9cfcec58b26

Observation 7a147ee5-e9af-423c-bcb8-e23668805dce · outbound

This paper cites Treerpo: Tree relative policy optimization, 2025 b.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Treerpo: Tree relative policy optimization, 2025 b

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.699710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.699710Z digest=sha256:5b48a6b82bdc5d800f63029a7a895f103c2e8192f3e7bbaa62f61d0fd7185efd

Observation cca10bc2-ba30-4015-879a-2d47cc5917f3 · outbound

This paper cites LIMO: Less is More for Reasoning.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance LIMO: Less is More for Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.703635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.703635Z digest=sha256:299200868050deed80f179b789ca16b203016bb0f2a11e5690303be9080c5eb4

Observation 6909db26-c4a7-4eba-b806-47763d3c1d22 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Dapo: An open-source llm reinforcement learning system at scale

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:47.917925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.707418Z digest=sha256:bd727ac6055ee0937ccdef6247211fd93f1829aab385651d224b8f99a6ee9c18

Observation c7e23046-74c3-432f-bca3-3e320e6856b8 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.711017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.711017Z digest=sha256:60483867b829e98c69058781d533e9b1852755e84344e66810229330d1a54305

Observation 77fd2a2b-5bef-42fc-aafd-5b6f5da33df2 · outbound

This paper cites R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:19:47.905804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.714963Z digest=sha256:4bb9a463dfa46344f1b46e4642a855c5a2ea51b5117e0dac749567aa56752b36

Observation d084dc3c-910a-4cbc-9535-8d8786694e82 · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.719051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.719051Z digest=sha256:ac8d04a12993dc2fab7dc7489e3391bdeeac884c896cf30f2d46634529e52ace

Observation bfc47b08-9b7a-4a03-80dd-de0235f3c273 · outbound

This paper cites Disco balances the scales: Adaptive domain- and difficulty-aware reinforcement learning on imbalanced data, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Disco balances the scales: Adaptive domain- and difficulty-aware reinforcement learning on imbalanced data, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.723728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.723728Z digest=sha256:7900765bb57d99e0895f035aa0416a14a411d08810bcc12bf75cf20f8a6a46d8

Observation a779c09c-06eb-4a63-97fe-ad541edafe3d · outbound

This paper cites A technical study into 0.5b reasoning language models, 2025.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A technical study into 0.5b reasoning language models, 2025

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-15T17:19:46.821203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:19:46.727588Z digest=sha256:a973e85480266c240b1b7ca9855d7b7a4d63c0d1ac96bfcbba8ebd964dc3e0ac

Pith citing papers

Observation c77f7630-e7d6-44fb-b63a-597701da933d · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:05:31.574588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:a0d159d04c0029a01b1920588c24f599cae8a4bc9c752e014a8ebcc758126033

Observation 7199d10d-8201-4edd-bc17-a5419835fcc6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:48.905173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:6ff0cbbd2306207ce8a367d00d7237a4c5f880d07cfa35e1021522d30535e17c

Observation 965b4760-9fc6-4212-9403-9f045c9e2cf2 · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:28:52.857138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T18:28:06.253200Z digest=sha256:e1db609d058de73204f44737a73c8b456a49a2c50cddfc4a0474593f492cec5d

Observation 352b23a7-b5db-4e71-9ddd-108c87290dcf · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:00.970934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T19:37:18.563704Z digest=sha256:62fa414dc9264d824159a50e5cf7c8003cae5ef35db5febaf6b1f0a71edc20ff

Observation a7c748f7-a37a-483c-9dcc-c030f20fa417 · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:32.886708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:32.886708Z digest=sha256:e1eb65c214ebdd7936c14dcd86eb4fa1120f44792737b684b6f694acd928a2f6