Pith. sign in

Paper Citation Record · LEDGER

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 18 inbound Pith citation observations for arXiv:2507.21848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21848 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:26:11.192134Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:40.838440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.935150Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d77ac373-6e99-49ab-a72e-7a0cf74997cd · outbound

This paper cites Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.159006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.159006Z digest=sha256:6b414e8a4bb1b715f442b058f848cf6586b97dc94f6e3fe578b2ff1a467bdb09

Observation d0b5dd49-9c09-4373-a9ad-b18ec0b14f2e · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Reasoning with Exploration: An Entropy Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.126635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.126635Z digest=sha256:d86165ad41a94e5462a5410a45232308ceb65c49b59354cb60f1be59da670d54

Observation 2f9b77b4-4d73-438f-ae59-802b873f62e6 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.130126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.130126Z digest=sha256:a23e111d5a85d73d47383fb5bffbad9acb52b23e5ee18fac91340df64c22a9b8

Observation ae609b57-5ef6-44cf-8f75-b48d85a05bab · outbound

This paper cites arXiv preprint arXiv:2504.05185.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2504.05185

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.133984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.133984Z digest=sha256:cdcab3eba8948371beda6a53c91874b4c34f54a86f2936ed0fbbccc5655be490

Observation 43d43ab2-3aa5-49da-8317-3a3b660ecc44 · outbound

This paper cites One-shot Entropy Minimization.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity One-shot Entropy Minimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.136900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.136900Z digest=sha256:e82f8bed584aaee43fbdf337c2d34ee93753525be8f008ce77b3e1068ef4430d

Observation ccdf164c-bf13-48e9-943b-2c4477f3f8e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.140006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.140006Z digest=sha256:8831f860f59343732ac955631ba442ffbe0f8e0db5e062f7f4e87ee4e90e5b6a

Observation 1f60f95d-8c20-4a20-b7b4-c6434b7f16bb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.149829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.149829Z digest=sha256:2683527d2f7d0962528ac4c36782d2fe488cdc30a2abc7362135559a064ef3e9

Observation 636b9b26-0998-42ca-942d-608d8a23df06 · outbound

This paper cites OpenAI o1 System Card.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.152824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.152824Z digest=sha256:bb7b52e02160b9af844d59c920bfebb73924a095c0f5a187088afb8f7c0fca8a

Observation 43b0583d-889d-4377-9a2d-eec2da785756 · outbound

This paper cites s1: Simple test-time scaling.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.161943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.161943Z digest=sha256:1df0b53484686f1e6d1622b4c8e076ad3712a13baf5c6b03129baf2f457664bd

Observation e8f6cb66-e737-4fd9-ae59-baf9f5a0791c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.167724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.167724Z digest=sha256:f5fae3dd6593e8a7e1d64d1c83a5b5e18c839ef97c00d98afed21538ca21bbe5

Observation d869d484-ae55-4538-8f5e-cefedf7cf522 · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.170585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.170585Z digest=sha256:e079d10b6b665beee12d72061f8f4cd29e5332b52817327017474099d8a44965

Observation 10859b75-3b04-4d6a-8f61-5a3f70f0b845 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.173550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.173550Z digest=sha256:1ef81a63d2ba233321ebe1c6b0152086fac413a240c210d20285c4f1e99f31d1

Observation a5081309-1f68-478d-bee0-953e78a9b6a2 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.177203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.177203Z digest=sha256:2109d849983574088786e134feefaca192b260abf0256bae3f857f1b68f88007

Observation 67b691d7-64df-43f5-9953-dbe20f528a52 · outbound

This paper cites arXiv preprint arXiv:2506.01713.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2506.01713

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.180394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.180394Z digest=sha256:65ca80d3d84023e0509a913a3088abd57f2725b5dea62031471ffa67a54fc3aa

Observation 35a5df91-f167-48b6-8b40-cd7dc059882d · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.183062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.183062Z digest=sha256:004a7f03bd403fa05ac7fd3a3f4dc113278d34ff256257e7916210cf303909e3

Observation 7fe85726-aa59-461a-a040-c4053a797ed2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.186011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.186011Z digest=sha256:6450767a1014f40e7d521dfa3c782633e13b595fc2329c32168596ec0baf57b9

Observation 79f6c4c2-a856-43d0-a4a3-0617e3b1b1ab · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.188931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.188931Z digest=sha256:24ae005d26aceb86da5a0854805a9aadb91093f00e42486fc8bb0757f9f1b47a

Observation 1c632aa7-4cc9-45bc-b549-9787c239eba3 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.192134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.192134Z digest=sha256:83c04dd2b6e741a37b307b093b75ba6ab00e6c8e29648e02635b843f07017719

Observation 215984af-1c66-40ea-892b-a9c8e23202bf · outbound

This paper cites Proximal Policy Optimization Algorithms.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.164984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.164984Z digest=sha256:22ee6b9655474e9f7ed2d686a3940b78b4fd1913a456a71da0630721290dab08

Observation 6a781e4e-ebb1-4424-b8b4-7cee598c08fd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.146760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.146760Z digest=sha256:fe4f15965d4759d3d301786f18840e8bbc05ae9d60d0e435d17ef2d06113ac2b

Observation 83f14dab-ca0d-4b49-8f7d-1a69efff6c58 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.155945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.155945Z digest=sha256:d18be33dfdd5e62ef04aa5d4842ae7819fe25da352ad864cdb140b0b65f09818

Observation db9b1184-dcf1-439e-b20d-5f4c0624b0df · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.143506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.143506Z digest=sha256:f874fd4ad2df31eaeb5ea088638e40bca390101d73db51f6a4bc717c8d8d01dc

Observation 69b01b73-13c2-4459-890e-334ee896733b · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.122616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.122616Z digest=sha256:d116510bce8b0b613c8ca4949bb71edc359c23f12c668572463843b3d120ad09

Pith citing papers

Observation cf2cc2bc-117b-4cd6-847a-ba838efe361b · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.764302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6f041df7a06be4b3d8e64db4b1498e65546fca28eded7a95a280a1bc240d4562

Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.634858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.634858Z digest=sha256:b0453ccb9b0c24410603a9e220037de5f1712b88e65fea199614ddaa5b8f9de5

Observation 728510a6-8d25-44f4-9359-2b8725dd0455 · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.435478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:69c75df841b83d34708f4f5f8389692f51040a98472eb78bccaeb0bd3a7342f2

Observation 6766ed2b-4d7d-4775-9f0d-4641884c93d1 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.674800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:92ddc3a2e155467584f4db72c5b38e19ff3474cbd0d47581bf84f5ad67da51ff

Observation 558a877a-ae86-41e9-b42b-c1f3052f6287 · inbound

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning cites this paper.

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T20:44:47.413838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:44:47.413838Z digest=sha256:006700eac6a4137b42d6144e081796401b02fbe95c8240a79feaadefe6c10740

Observation 9622d80b-8959-4571-84e3-629ea3086da1 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.951337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:a544e0e9da1c6ed2e99ba1dca44d4628b9820bff702cff2b09fe631fd3aa0cdb

Observation 4c9c248a-4723-453d-942f-2d66754ef541 · inbound

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance cites this paper.

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.056711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T08:05:36.206946Z digest=sha256:6d89b1fbefbc7812487c0827603b4650e58132b6b55d3e65b218a28b9f271bf2

Observation 58666f63-cb06-4f1a-9b0a-b2b0612dda90 · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.309046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T14:52:33.141693Z digest=sha256:d7851e691ab2a6b20add27b42d5a33fab3351d1f2f142d65d9d3a2a15e86889d

Observation 89914b81-2235-4847-bb65-3a9f4ecd70db · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.406585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:a581dcc3673737d9306bcf137880015c76358c839118d0c6fce2ca802ed2d4e1

Observation b88c5685-4fd8-4dd7-9e91-768ecff1be67 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.261365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-21T06:08:55.571828Z digest=sha256:08a211dadc659c58ad43d52b228de90b677250244056933f2fa8f7059044f701

Observation f711196d-2267-4ebf-9db1-99794f9289d3 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.310221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-30T17:15:53.286925Z digest=sha256:138ac8653dc55c1c7bc685dd5958e4b173d2169ecdcc564dea562dc10e6ada9e

Observation d573ebba-a89f-4ef3-9e76-5f893d5e3882 · inbound

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO cites this paper.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:49.989463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T23:54:32.621093Z digest=sha256:f24006988367a18be5092988c8cc91fcedd2c818b58c69ab24414bd2037a1e36

Observation 1d0ebd72-b3e4-4607-a14a-536159c26a1a · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.936872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:79868d5cc59f2b636d9158d33dcffcf2c28df085539852aff393dbef37bed35e

Observation e947bb10-d5b3-4419-b4d6-3f8825599d84 · inbound

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning cites this paper.

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:19.253487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T06:10:41.786328Z digest=sha256:c902f4b78bb87bf48bf27a91050eb0bd3f586e27dd78d5ca31d320589dc474cf

Observation 71968d98-9133-4017-8a89-4e45d84158cd · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.413060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:fafd50fa69409a415c3eef5d5cc09155cd751b64339e63a2e701041da92e3ae9

Observation 0a8f6b9b-0cad-4138-b6d3-a4bc95a6b36a · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T18:43:17.536451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:43:17.536451Z digest=sha256:5bbfcd0a3acb4ec2a2643656056077096f7217f568816f492cec4e4548751891

Observation bbb13f40-929a-4252-ab59-dd58226c2ec5 · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:47.185967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:52:47.185967Z digest=sha256:d4763c2e03f3c5d76406dcda912f327fd811647563e8eeced5024c50d10df81b

Observation 3b81bcac-b01f-4ce5-917d-27972528646d · inbound

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models cites this paper.

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:40.838440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:40.838440Z digest=sha256:f863ab07c2a322d9825b4e0d9a1708ffddfef31157b9eff88f72f05fa702d412