Pith. sign in

Paper Citation Record · LEDGER

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 35 inbound Pith citation observations for arXiv:2505.12346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12346 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:13.165308Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.427933Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ee99a11-40e6-45e4-bc87-f4d4f5fe37d7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.011959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.011959Z digest=sha256:5bfc997360f3eea6fb4a341603ee58f6c81e6d4f9026215e1619eecc28c0e35a

Observation 93197f6b-9e95-4cab-a620-cd82fcce7263 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.016281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.016281Z digest=sha256:48422a237c09c58e249220fd3d5257a2422ea58b3226b0be3fa7cfa13dbd6d3f

Observation 806637a3-f613-4a80-90d1-8ea4896e31ef · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.020339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.020339Z digest=sha256:827bda4654ef1093f6eec2864c81e288234d67c12654f3c54af0256f8c1a2f1a

Observation d72c7687-aa36-4f31-9f98-c9dde95e9424 · outbound

This paper cites GPT-4 Technical Report.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.024490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.024490Z digest=sha256:4badb3387ac3ad3fd81572af9724a8f6df9911c10e07921eb77f56ba06ad15ad

Observation 060360f2-1a43-4706-b4f0-d4dbcd2858d1 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.028343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.028343Z digest=sha256:94bb8907514a439d36dd92b4d9c9c5340d8ef3ace026bc9f63495e058ba5e96b

Observation 2502b848-428b-4b18-bfe3-feaa41f62102 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.032436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.032436Z digest=sha256:c9581be511ff5584d310478ee028ed2d177055e795fcba0c953a9acb7ef3fc0f

Observation f557445a-6657-4e71-8261-a5188b2cbcd7 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.036887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.036887Z digest=sha256:bcfb60e788978bc7f0566b514729cb8853300f7eb1d6ab2aef38131287f24feb

Observation 7825dc23-4a1e-4b4f-85c6-2153f653c563 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.040337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.040337Z digest=sha256:a697cec7d64490e18e506e7bf8d8ec40296269ebdfdd6b2d9dca2a73ca6769b4

Observation 788c11cf-4a69-464b-a656-4f9e647dda6b · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.044129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.044129Z digest=sha256:bda0923ac5ce6c0fa69ced1ceb975e1fc779e79d4a4563912a8766076b1beffc

Observation 2c595027-37fa-465e-83eb-0a25c13f14b6 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.047975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.047975Z digest=sha256:327155b1cf4e8dc903883a8b6facdcf87fcef55507756df81f4f6d84e78eb053

Observation 8f197b82-6299-4f5e-ba2f-a2d960eaec36 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.051993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.051993Z digest=sha256:4fdfa4564a9ef55abd6b065dae9c456fdcacd2be9334469e55c4ca4cb4a0ca65

Observation 3e992d9d-dbcf-4ad6-ae0e-c713228e4c24 · outbound

This paper cites OpenAI o1 System Card.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.056166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.056166Z digest=sha256:b23e6a6d52ad70ce6f1194feddb26a34641bcfd62d6123619c88654a3bfec116

Observation d74b4e0c-42e0-49e2-a291-df256a281dc7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.059812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.059812Z digest=sha256:92acf07d4831961874ef7bf5e3bc7652cc5f4348c36a36677bd1a15a1645a4a6

Observation 26c306b2-9fc6-4406-96a2-3acc9fc7fdf8 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization The claude 3 model family: Opus, sonnet, haiku

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.847921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.063624Z digest=sha256:5304aa69b35580b3464260e7178a66faa66d1b6d3323609b69def17610eea863

Observation 5677145c-b958-48a5-8a0f-6cfe2a052da7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.067375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.067375Z digest=sha256:f943c01218c1af2b9547945e456a8b81651e40f0380813da3cac2bf103887b0f

Observation 17b33c8d-64ff-47c5-883b-a9402670cbb0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.071250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.071250Z digest=sha256:3126fa019472677d70ffd90f0ef3540cd8424548c03a54a8eb57d36d97e3b8d4

Observation c9ac5c6f-f0bf-4be6-8bb9-114314a10633 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.074768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.074768Z digest=sha256:1ed1f6d999be22fb159a4539714baeae3eb73e8a3b644764f34e028e2e01a767

Observation 4d41709b-58e5-4e78-9b78-03a7a7b7d1a2 · outbound

This paper cites Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.078247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.078247Z digest=sha256:bc267723c575dbe44ee56d0c1c99d077d6bb78446b8ef5e52e59f2f499341151

Observation e58237ea-71b5-4bb8-b946-dbe383eb6477 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.081592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.081592Z digest=sha256:abfff50a36b5d1a4c608e73252db92c793b08888d24bd765832fb8dc82f718eb

Observation 411aa082-0154-4bea-993d-b0c547b36a80 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.085686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.085686Z digest=sha256:7bdae73e5ba7c73a3eff39dea3cdb3637409a871c47b613cc2048768425aea5d

Observation 9da99666-9fcb-4d92-a774-359878e57ad3 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.829638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.089360Z digest=sha256:1e84c68b3c9d9553f9dacce21023fdf6317cf56f2866d9cc514e2c8fd219c77e

Observation 62ca4571-c604-40a6-98ad-08bf2d5c3b01 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R Narasimhan.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Griffiths, Yuan Cao, and Karthik R Narasimhan

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.817290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.092868Z digest=sha256:32227a9d496f9b0f02cbeb518d3d384cfde907a943cfbc2039523ef36c590168

Observation 4de2cb60-2685-4eb1-94b8-529d495a03cc · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.805391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.096326Z digest=sha256:78a9b11aa2c365b39e75fae0ff8c8ec3c300f6092cc0ca4d24aaf06b0b750e77

Observation fe917928-6605-40af-9bea-2bc25baec380 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.099891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.099891Z digest=sha256:b9427d18833ee37c2b65d2bac2657c5d06535f8d3bf97d56366f4e262e04e39e

Observation 1cb3a471-1774-4bd8-968a-38b86e82858b · outbound

This paper cites Curriculum learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Curriculum learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.103957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.103957Z digest=sha256:91c4d7945a3f48a772dab93d4595102321b67281787a6ba45cba62f73cf3c4bb

Observation 656e5340-d350-417b-b530-0b29a7b9794d · outbound

This paper cites A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.107434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.107434Z digest=sha256:513f35ebd00800d37c207881bf7d7d7fc657dfbcc25949c1d41a34a942982c97

Observation 0597bbb4-e21c-4d68-a9f7-20826ecbbe40 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.111000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.111000Z digest=sha256:cdd3fcd51cc8d90d4697038cbfbb8181e937a358a2eaefa613a701ec04a82226

Observation 76b47959-eda1-4b51-91e6-1a27791f76c6 · outbound

This paper cites LIMO: Less is More for Reasoning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization LIMO: Less is More for Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.115113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.115113Z digest=sha256:c8f595cf8df472f16f6a2b54676cffb21042665783fd208c8b2c339ca5a1f415

Observation 80079864-e5b5-42e5-8828-cfe4073c139c · outbound

This paper cites ReST- MCTS*: LLM self-training via process reward guided tree search.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization ReST- MCTS*: LLM self-training via process reward guided tree search

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.771843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.118997Z digest=sha256:f7702d42125f7ef30d91bfb0504262415570130e53e86f1f304ac3dce318fb36

Observation be79bc08-9e40-47b3-818d-42df74bdf59f · outbound

This paper cites Proximal Policy Optimization Algorithms.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.122315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.122315Z digest=sha256:b6794a1da91ef9cb352de42db8ddcf503429f736a6409b6161199f1dc4250cb9

Observation 53cb17cb-c18c-4078-8e31-ddab9477267e · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.125943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.125943Z digest=sha256:e8e22b0ed6f667dec874187df980df097b3d513bb732e95bcc77ace97cfe8afb

Observation 3cc60ea2-ddc0-474e-9acc-a62709c82019 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Measuring mathematical problem solving with the MATH dataset

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.751927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.129548Z digest=sha256:2ed5ef3f05291e78ab8a7d62397d7c785116f9f2db8d1b13c4c938b90b4cc146

Observation 80a5c625-2dd1-43d7-bb23-158363ec38c9 · outbound

This paper cites Solving quantitative reasoning problems with language models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Solving quantitative reasoning problems with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.740603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.132960Z digest=sha256:4125e166a4dc72b831450b176feb95482b52d892bd3223f2117521dfed65a959

Observation 3c608ec6-defa-464d-9b97-74487315f1df · outbound

This paper cites Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.728314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:40:13.136266Z digest=sha256:df002b4a347f483b82172e4dd19ebaac20d4e853f5d9ac6d8ec390d19c558509

Observation ae2d74c8-0d7e-42d9-83b0-244a2a88dffc · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.139709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.139709Z digest=sha256:8e37edd4a092c18464822cf87bd791b1244422a0eec8f5dfc68b18857cf8ce4c

Observation 879da7f4-6451-4c9d-8d98-f5e0fac47948 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.143452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.143452Z digest=sha256:30bff222404d72b1aef1731859f851a5407f2d1f2718006b11afd3f97a2fdc5c

Observation f2527d3e-6789-464e-b0c6-c7f67fa7a817 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.147339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.147339Z digest=sha256:8e8b86ec1d0ec8f65ef09233ed83d7fd9eddb772527665be8cbe14dfb3c59a69

Observation e18e2a4c-53a7-4fbb-91f4-f5ad7d9d0b9e · outbound

This paper cites Qwen2.5 Technical Report.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.151066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.151066Z digest=sha256:93b3a06fcfed91c464eeb44166aa367ef53861415e89004775dbbbccbdca9698

Observation 00c90c4c-6483-41a1-bd50-b4ab1ce0c251 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Evaluating Large Language Models Trained on Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.154555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.154555Z digest=sha256:5e934856e177f47cd47c74fa2638e454dd83e191c26144d269c53b190dd7471e

Observation ee56f8a5-890a-45f2-9dae-01ef05def261 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.157996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.157996Z digest=sha256:e310766a79e70226d332e0cdb0b121964be7968c1719c9cb55cc86da60978845

Observation c22926a8-31ba-4a2b-9709-4450195aa258 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.161587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.161587Z digest=sha256:da5dc7d3d8545751e40c5ad17b40b38026b7c27f63bd92ec411c427a8b95beeb

Observation 69e40fbb-3f69-438d-b86e-5541bab89966 · outbound

This paper cites Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.165308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.165308Z digest=sha256:a8d5102ae44a3aa7f4659196eff777bc3136e5d602ccb30af5a3e8efb5993d3c

Pith citing papers

Observation 7b4aba80-d73a-4225-914e-cd2b8ce53e20 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.190500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.190500Z digest=sha256:b97bbf1babff88adb04d363fb2ef9bd9c1755759824d18380546f5c55668c5b8

Observation 0b9f6000-6e9b-48f9-87c5-38a7c22a23d0 · inbound

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation cites this paper.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.170825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.170825Z digest=sha256:07e499c2db4f84b30713f7e1bdff3692c355cc06e37aa417e49377f8ba8539b7

Observation 51f79782-9ec8-4f33-83a3-77ede35621bd · inbound

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework cites this paper.

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:59:40.639865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:59:40.639865Z digest=sha256:c3cd59aa14fd3f3042399f2051190c7d6fc96fd2b51fad4a13b091764efa3776

Observation 69b01b73-13c2-4459-890e-334ee896733b · inbound

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity cites this paper.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.122616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.122616Z digest=sha256:c6063a8374749574749509960eaf06e84030cfa713a6baf130cf760c35f2c872

Observation 892fbaba-8018-4613-907c-bbe0cdaaa6cc · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.452911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.452911Z digest=sha256:33d2e70f26f0fafe25be96a518c61e770c213f97e7e36d26d00c574e72364d87

Observation 1d7da6ce-8da8-4a9a-8688-18fe4b8d55e8 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.084591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:0eb460437b2a4a4eb108a406c07c0896bd106a9f73d00de17eea6be5f8dd8853

Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.258016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.258016Z digest=sha256:66f3fbc85eee3479ffc19e9e8e0eb4cf616ecb30fd10a5500165225c688869b5

Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.055373Z digest=sha256:e454520c5b3d45113bcd6e795873db6aaef6a9b7bf1f6ac7bae6e865b1978887

Observation 9d457360-c1b1-41e6-8da8-77d54a0f8373 · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.407632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:a5223be8947b7588250a56cec163379dc3f747ddcd50ba24730e4006e1bba588

Observation f273ed6b-ec28-433b-b6e8-b982d2353ee8 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.768253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.768253Z digest=sha256:e1098c7f37d324492a23888561ce00cf6238739078a282e654471ba8a83060fb

Observation 4b031acc-1e60-4a7e-9592-a5d988094944 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:03.940645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:03.940645Z digest=sha256:6cde7916c28ebf9fdada45ce16ab0eca6875ed522685c0b42bac59a40efec6a6

Observation da19f184-186d-4aed-b54c-74ac41b3cf70 · inbound

Self-Distilled RLVR cites this paper.

Self-Distilled RLVR SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.563241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T19:43:46.623267Z digest=sha256:04ed0ba90f3efb8e15d636c5d601fa9e455fa8c2484bdf01717d839c67560138

Observation 97aa497d-3246-4818-8110-ce4dfb077a56 · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.244538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:8a45a25df68e4a2bcd8997d2cbc679a32fe146c638eb6650843533ede7a4f468

Observation caa9ebf4-db47-47db-b183-b28db1c89bbc · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:55.199930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:715aab7c358edf1d4f3969895c5e3f611b593a7f7d6be5aba5a3caa58c9e1cb7

Observation 40b4c59d-42a6-46d5-a2f3-85242df039df · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:41:26.300159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:879c797142b116f5cdd3bb9046fa79dcbbc768188f71eae2e729f9dc0fa09e45

Observation 131bf09f-eb00-4823-b093-0ab1106f3131 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.128262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:b9f71924dd57e8d160e4ba55ce5b85b27247a01f7360c8bab144883f6dbd42b9

Observation c9ae2cc9-c05b-4afd-b031-5b719c166158 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:48.864199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:31c607ee8aa103477d3ec9fe6e31129057575aedc4c3eb92d2668e2014552df0

Observation 7ddcc6ed-9c0a-4a53-b513-ecf095e389f3 · inbound

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers cites this paper.

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:09.550349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T17:11:31.826233Z digest=sha256:f4709590ae39c11aad6c2c12f06ad5e05011d580fce2dfcc83f579468a938f57

Observation 6c96d400-c4e2-4a8f-8d3c-f44cd106e081 · inbound

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models cites this paper.

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.629634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T00:48:47.213681Z digest=sha256:681c2a81cec6fa980487c5d1b373b6f278f0ffa795fe4d27468a44d74b2218f3

Observation 7f265df0-44b6-4a12-a82c-653cfcab9e64 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:29.447553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:5d44a826284e2f18f478ee291ac8ce1ff68ab6ba7c394ae9fe785a35a181837b

Observation da4eb23b-efe3-4d94-bfdc-eed8427c032c · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.790348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:f6007cb8cfea7fd59bcdfbe45ebc773e24b6ca32383c670fdf5b63320f449687

Observation 94c188f2-641f-4ea3-8252-6faeb9c094f9 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.115200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:0448882df57229295712dfc9b5fd7029bcf13904312c5c171b5f738c964fbc94

Observation 0a337026-3afb-47eb-956a-f88f1dcfd852 · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:43:05.721502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:7c9d00d8ab75b854bad2d70ffce83746228677d6d405b558975efe2a62f10057

Observation 82fbf2f1-0959-46e2-a70b-f94cd95608b1 · inbound

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization cites this paper.

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.757498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T08:53:02.051453Z digest=sha256:60eb4049da908d8fb0dc33bd1d4537084091584449277dce99f52abc1b793c86

Observation 45242034-0215-4725-8b8e-0e4ae85a8d30 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.407524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:05db45d0497dabe627017942949eb1e2ec9c7cbbcc2555f202c6725d309086e7

Observation 7416aa03-8007-4fb4-a17f-6c2dc8e53bbf · inbound

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data cites this paper.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.608391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:2d8d18e5aadb0d28bf87ef8e6971da4c4538aab9759512dc74b3bdfba557e4d6

Observation 2383a679-d435-4953-9f81-f7ece23e5234 · inbound

Self-Distilled Policy Gradient cites this paper.

Self-Distilled Policy Gradient SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.906762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T11:24:11.878439Z digest=sha256:2739496e252375fd4abffc9c60c4cbaa90e8f19212813d60331fcd0d4691ccac

Observation d12d0f00-9328-4da1-be46-364c0af13427 · inbound

Reinforcement Learning from Rich Feedback with Distributional DAgger cites this paper.

Reinforcement Learning from Rich Feedback with Distributional DAgger SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:45.915273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T06:44:56.667364Z digest=sha256:63a2c81f30f6113d3cb4f89bef0a3886cc571fa08ff00d293252408a2dfd4b80

Observation d8439ce1-be73-4c88-a3ba-b2bbd9d8e9e5 · inbound

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models cites this paper.

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.585403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T18:31:21.493677Z digest=sha256:9fb631435bdd28f76034c3f97b01dbf4f7263a63f4b211fd5c5b177304d4b1a8

Observation 40290d98-3b46-434d-8d96-eb9c56da9001 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.902366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:737c82deadafe56c96cfffc465bd4a67853ac1baebd68551d80cd1f6967c3f28

Observation eaa9e059-daff-422f-8897-f28c490542eb · inbound

Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System cites this paper.

Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.837064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T21:43:01.949350Z digest=sha256:a299e5503793b06ae66c9760125d57b53c6916b66c816617e59feee1df5e7f9a

Observation 14275d2a-83f5-485b-b2f0-ea9c0e1c32d6 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.560962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:e4d21a1a29bb8e5b406171bf0533aaa293fc96f08e9008381fb61dbeea064c87

Observation 8829c247-63e7-43ce-b82f-14c140911edd · inbound

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning cites this paper.

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.429607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T04:49:28.598430Z digest=sha256:f940acb250bc5250b50398094031318b210cdcc0a56f21ea8a3fd9cc0ab9f143

Observation 00ded8ff-7d33-420b-a0fd-003270913020 · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.387378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:c78dd2877fa37acfdffbcd47ee431d9accf9a64ad9ee9d26bdc998bfcf133a77

Observation 7c8263bc-0673-4a8f-83b1-d2cf75e97c99 · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:42.043252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:c828a6ae65d7b18563a6d039452b5ba4ae871a4e74faae94e3d6589901b3fd94