Pith. sign in

Paper Citation Record · LEDGER

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 35 inbound Pith citation observations for arXiv:2505.12346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12346 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:13.165308Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.427933Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ee99a11-40e6-45e4-bc87-f4d4f5fe37d7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.011959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.011959Z digest=sha256:5bfc997360f3eea6fb4a341603ee58f6c81e6d4f9026215e1619eecc28c0e35a

Observation 93197f6b-9e95-4cab-a620-cd82fcce7263 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.016281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.016281Z digest=sha256:48422a237c09c58e249220fd3d5257a2422ea58b3226b0be3fa7cfa13dbd6d3f

Observation 806637a3-f613-4a80-90d1-8ea4896e31ef · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.020339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.020339Z digest=sha256:827bda4654ef1093f6eec2864c81e288234d67c12654f3c54af0256f8c1a2f1a

Observation d72c7687-aa36-4f31-9f98-c9dde95e9424 · outbound

This paper cites GPT-4 Technical Report.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.024490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.024490Z digest=sha256:f66f9ae601193e27a4ecd95e52c3d539802f402e62f856c5d112f2addaf5feee

Observation 060360f2-1a43-4706-b4f0-d4dbcd2858d1 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.028343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.028343Z digest=sha256:94bb8907514a439d36dd92b4d9c9c5340d8ef3ace026bc9f63495e058ba5e96b

Observation 2502b848-428b-4b18-bfe3-feaa41f62102 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.032436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.032436Z digest=sha256:c9581be511ff5584d310478ee028ed2d177055e795fcba0c953a9acb7ef3fc0f

Observation f557445a-6657-4e71-8261-a5188b2cbcd7 · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.036887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.036887Z digest=sha256:bcfb60e788978bc7f0566b514729cb8853300f7eb1d6ab2aef38131287f24feb

Observation 7825dc23-4a1e-4b4f-85c6-2153f653c563 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.040337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.040337Z digest=sha256:a697cec7d64490e18e506e7bf8d8ec40296269ebdfdd6b2d9dca2a73ca6769b4

Observation 788c11cf-4a69-464b-a656-4f9e647dda6b · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.044129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.044129Z digest=sha256:bda0923ac5ce6c0fa69ced1ceb975e1fc779e79d4a4563912a8766076b1beffc

Observation 2c595027-37fa-465e-83eb-0a25c13f14b6 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.047975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.047975Z digest=sha256:327155b1cf4e8dc903883a8b6facdcf87fcef55507756df81f4f6d84e78eb053

Observation 8f197b82-6299-4f5e-ba2f-a2d960eaec36 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.051993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.051993Z digest=sha256:4fdfa4564a9ef55abd6b065dae9c456fdcacd2be9334469e55c4ca4cb4a0ca65

Observation 3e992d9d-dbcf-4ad6-ae0e-c713228e4c24 · outbound

This paper cites OpenAI o1 System Card.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.056166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.056166Z digest=sha256:b23e6a6d52ad70ce6f1194feddb26a34641bcfd62d6123619c88654a3bfec116

Observation d74b4e0c-42e0-49e2-a291-df256a281dc7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.059812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.059812Z digest=sha256:92acf07d4831961874ef7bf5e3bc7652cc5f4348c36a36677bd1a15a1645a4a6

Observation 26c306b2-9fc6-4406-96a2-3acc9fc7fdf8 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization The claude 3 model family: Opus, sonnet, haiku

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.847921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.063624Z digest=sha256:0021fa2913c8e19b28d8ba45abd8cb5c65c7258d4849d377eed50e18629454dd

Observation 5677145c-b958-48a5-8a0f-6cfe2a052da7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.067375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.067375Z digest=sha256:f943c01218c1af2b9547945e456a8b81651e40f0380813da3cac2bf103887b0f

Observation 17b33c8d-64ff-47c5-883b-a9402670cbb0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.071250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.071250Z digest=sha256:3126fa019472677d70ffd90f0ef3540cd8424548c03a54a8eb57d36d97e3b8d4

Observation c9ac5c6f-f0bf-4be6-8bb9-114314a10633 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.074768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.074768Z digest=sha256:1ed1f6d999be22fb159a4539714baeae3eb73e8a3b644764f34e028e2e01a767

Observation 4d41709b-58e5-4e78-9b78-03a7a7b7d1a2 · outbound

This paper cites Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.078247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.078247Z digest=sha256:bc267723c575dbe44ee56d0c1c99d077d6bb78446b8ef5e52e59f2f499341151

Observation e58237ea-71b5-4bb8-b946-dbe383eb6477 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.081592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.081592Z digest=sha256:abfff50a36b5d1a4c608e73252db92c793b08888d24bd765832fb8dc82f718eb

Observation 411aa082-0154-4bea-993d-b0c547b36a80 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.085686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.085686Z digest=sha256:7bdae73e5ba7c73a3eff39dea3cdb3637409a871c47b613cc2048768425aea5d

Observation 9da99666-9fcb-4d92-a774-359878e57ad3 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.829638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.089360Z digest=sha256:d660f76ad68eaba159b9175e608f1ace934e752559c9b73c4254f6686f627968

Observation 62ca4571-c604-40a6-98ad-08bf2d5c3b01 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik R Narasimhan.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Griffiths, Yuan Cao, and Karthik R Narasimhan

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.817290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.092868Z digest=sha256:7b4a096f7ae6648cc9fe615d72decfb9ff0496c7025ebc78bc1789a8230392a9

Observation 4de2cb60-2685-4eb1-94b8-529d495a03cc · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.805391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.096326Z digest=sha256:6576d60110a253c6e538162fedabdff30d438bcddb4caa2b0c0dbcb303b9482c

Observation fe917928-6605-40af-9bea-2bc25baec380 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.099891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.099891Z digest=sha256:b9427d18833ee37c2b65d2bac2657c5d06535f8d3bf97d56366f4e262e04e39e

Observation 1cb3a471-1774-4bd8-968a-38b86e82858b · outbound

This paper cites Curriculum learning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Curriculum learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.103957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.103957Z digest=sha256:91c4d7945a3f48a772dab93d4595102321b67281787a6ba45cba62f73cf3c4bb

Observation 656e5340-d350-417b-b530-0b29a7b9794d · outbound

This paper cites A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.107434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.107434Z digest=sha256:513f35ebd00800d37c207881bf7d7d7fc657dfbcc25949c1d41a34a942982c97

Observation 0597bbb4-e21c-4d68-a9f7-20826ecbbe40 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.111000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.111000Z digest=sha256:cdd3fcd51cc8d90d4697038cbfbb8181e937a358a2eaefa613a701ec04a82226

Observation 76b47959-eda1-4b51-91e6-1a27791f76c6 · outbound

This paper cites LIMO: Less is More for Reasoning.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization LIMO: Less is More for Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.115113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.115113Z digest=sha256:c8f595cf8df472f16f6a2b54676cffb21042665783fd208c8b2c339ca5a1f415

Observation 80079864-e5b5-42e5-8828-cfe4073c139c · outbound

This paper cites ReST- MCTS*: LLM self-training via process reward guided tree search.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization ReST- MCTS*: LLM self-training via process reward guided tree search

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.771843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.118997Z digest=sha256:986c5c21556a7f39b6626a379d1001df4e786c1748c16555c9f00ecbf810be66

Observation be79bc08-9e40-47b3-818d-42df74bdf59f · outbound

This paper cites Proximal Policy Optimization Algorithms.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.122315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.122315Z digest=sha256:b6794a1da91ef9cb352de42db8ddcf503429f736a6409b6161199f1dc4250cb9

Observation 53cb17cb-c18c-4078-8e31-ddab9477267e · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.125943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.125943Z digest=sha256:e8e22b0ed6f667dec874187df980df097b3d513bb732e95bcc77ace97cfe8afb

Observation 3cc60ea2-ddc0-474e-9acc-a62709c82019 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Measuring mathematical problem solving with the MATH dataset

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.751927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.129548Z digest=sha256:8a394f519a8733d0896271e0575c2b3a308989dbd0fa56f5604b99c9690fbb78

Observation 80a5c625-2dd1-43d7-bb23-158363ec38c9 · outbound

This paper cites Solving quantitative reasoning problems with language models.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Solving quantitative reasoning problems with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.740603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.132960Z digest=sha256:c5adf070934bc2b6dc005d410858ff530edea4f423a959c3fdc4f3e818958ac2

Observation 3c608ec6-defa-464d-9b97-74487315f1df · outbound

This paper cites Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:40:13.728314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:40:13.136266Z digest=sha256:dbb3f4b772305bc6a67a4574e197152614c39c0d5900e67cd1ec54043749a8df

Observation ae2d74c8-0d7e-42d9-83b0-244a2a88dffc · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.139709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.139709Z digest=sha256:8e37edd4a092c18464822cf87bd791b1244422a0eec8f5dfc68b18857cf8ce4c

Observation 879da7f4-6451-4c9d-8d98-f5e0fac47948 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.143452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.143452Z digest=sha256:30bff222404d72b1aef1731859f851a5407f2d1f2718006b11afd3f97a2fdc5c

Observation f2527d3e-6789-464e-b0c6-c7f67fa7a817 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.147339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.147339Z digest=sha256:8e8b86ec1d0ec8f65ef09233ed83d7fd9eddb772527665be8cbe14dfb3c59a69

Observation e18e2a4c-53a7-4fbb-91f4-f5ad7d9d0b9e · outbound

This paper cites Qwen2.5 Technical Report.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.151066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.151066Z digest=sha256:93b3a06fcfed91c464eeb44166aa367ef53861415e89004775dbbbccbdca9698

Observation 00c90c4c-6483-41a1-bd50-b4ab1ce0c251 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Evaluating Large Language Models Trained on Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.154555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.154555Z digest=sha256:5e934856e177f47cd47c74fa2638e454dd83e191c26144d269c53b190dd7471e

Observation ee56f8a5-890a-45f2-9dae-01ef05def261 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.157996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.157996Z digest=sha256:e310766a79e70226d332e0cdb0b121964be7968c1719c9cb55cc86da60978845

Observation c22926a8-31ba-4a2b-9709-4450195aa258 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.161587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.161587Z digest=sha256:da5dc7d3d8545751e40c5ad17b40b38026b7c27f63bd92ec411c427a8b95beeb

Observation 69e40fbb-3f69-438d-b86e-5541bab89966 · outbound

This paper cites Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.165308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.165308Z digest=sha256:a8d5102ae44a3aa7f4659196eff777bc3136e5d602ccb30af5a3e8efb5993d3c

Pith citing papers

Observation 7b4aba80-d73a-4225-914e-cd2b8ce53e20 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.190500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.190500Z digest=sha256:b97bbf1babff88adb04d363fb2ef9bd9c1755759824d18380546f5c55668c5b8

Observation 0b9f6000-6e9b-48f9-87c5-38a7c22a23d0 · inbound

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation cites this paper.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.170825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.170825Z digest=sha256:07e499c2db4f84b30713f7e1bdff3692c355cc06e37aa417e49377f8ba8539b7

Observation 51f79782-9ec8-4f33-83a3-77ede35621bd · inbound

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework cites this paper.

Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:59:40.639865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:59:40.639865Z digest=sha256:c3cd59aa14fd3f3042399f2051190c7d6fc96fd2b51fad4a13b091764efa3776

Observation 69b01b73-13c2-4459-890e-334ee896733b · inbound

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity cites this paper.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.122616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.122616Z digest=sha256:d116510bce8b0b613c8ca4949bb71edc359c23f12c668572463843b3d120ad09

Observation 892fbaba-8018-4613-907c-bbe0cdaaa6cc · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.452911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.452911Z digest=sha256:33d2e70f26f0fafe25be96a518c61e770c213f97e7e36d26d00c574e72364d87

Observation 1d7da6ce-8da8-4a9a-8688-18fe4b8d55e8 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.084591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:55227918779440ed8ecf60aa7d8cb0006df8b74787aec8e12a7fa6d65ea83e63

Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.258016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.258016Z digest=sha256:66f3fbc85eee3479ffc19e9e8e0eb4cf616ecb30fd10a5500165225c688869b5

Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.055373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.055373Z digest=sha256:e454520c5b3d45113bcd6e795873db6aaef6a9b7bf1f6ac7bae6e865b1978887

Observation 9d457360-c1b1-41e6-8da8-77d54a0f8373 · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.407632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:5705ae047143dfbb6f6be66a7830f9c1ab9637189e04e3a80df124909e954fe9

Observation f273ed6b-ec28-433b-b6e8-b982d2353ee8 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.768253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.768253Z digest=sha256:e1098c7f37d324492a23888561ce00cf6238739078a282e654471ba8a83060fb

Observation 4b031acc-1e60-4a7e-9592-a5d988094944 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:03.940645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:03.940645Z digest=sha256:6cde7916c28ebf9fdada45ce16ab0eca6875ed522685c0b42bac59a40efec6a6

Observation da19f184-186d-4aed-b54c-74ac41b3cf70 · inbound

Self-Distilled RLVR cites this paper.

Self-Distilled RLVR SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.563241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T19:43:46.623267Z digest=sha256:dff1d0cf2b3870888c7939f741e979dc301046c27d22188eef14913a57fcae89

Observation 97aa497d-3246-4818-8110-ce4dfb077a56 · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.244538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:fb5f09b3cebb4fa3acf67b382865b1b1ebf305c736058a9931a8513a5646eac1

Observation caa9ebf4-db47-47db-b183-b28db1c89bbc · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:55.199930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:1d32248536a92aa4f3bfdbe103af84e9d3ed447d67676e43e2810710796d1cc4

Observation 40b4c59d-42a6-46d5-a2f3-85242df039df · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:41:26.300159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:25fc854c5ec7d86a9067d395fb9d3f91ce40793f661dcd5fc351e11600c1157f

Observation 131bf09f-eb00-4823-b093-0ab1106f3131 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.128262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:e86c0d77d48f9c1d697576ad3dac77cf9b0650fbbcf80a8416fbaad777490234

Observation c9ae2cc9-c05b-4afd-b031-5b719c166158 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:48.864199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:78ddfb900c3021780440e0ca78fce14df3f682537b336e8c65e44173563f0ae4

Observation 7ddcc6ed-9c0a-4a53-b513-ecf095e389f3 · inbound

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers cites this paper.

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:09.550349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T17:11:31.826233Z digest=sha256:2a81433c19aa49d65eedb75f01076c0ba85f9e03d83b1f679470d6dd21a289d7

Observation 6c96d400-c4e2-4a8f-8d3c-f44cd106e081 · inbound

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models cites this paper.

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.629634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T00:48:47.213681Z digest=sha256:be99812bce56e1262358dd38a3cd4c4728e95b34cfd8da9dc2b7773d1981a1da

Observation 7f265df0-44b6-4a12-a82c-653cfcab9e64 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:29.447553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:caa9bff294fe2411a6ac1255f60c6a60d8c699f14ca8cfb0bb9daf8f30b6fd76

Observation da4eb23b-efe3-4d94-bfdc-eed8427c032c · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.790348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:82f74b97d790e1968678d4fa9d0d949accf690ef52653fd22d9fd876003a01d9

Observation 94c188f2-641f-4ea3-8252-6faeb9c094f9 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.115200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:301392955374b9dd5a067d7b408cac0fc450a5dc66319d51b19417a2e2dfad68

Observation 0a337026-3afb-47eb-956a-f88f1dcfd852 · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:43:05.721502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:1dd4660400b8e136a09cd51321e570827534656b89bdca3cbba610528b488117

Observation 82fbf2f1-0959-46e2-a70b-f94cd95608b1 · inbound

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization cites this paper.

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.757498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T08:53:02.051453Z digest=sha256:66eb937749ed070b0a10d28ace9afa3625547c6aa674141622f4c08ff168c7dc

Observation 45242034-0215-4725-8b8e-0e4ae85a8d30 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.407524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:3db3a4e87b0ee863518972f53c12d17c8b11345cfd0f2e12b84e46b08892945f

Observation 7416aa03-8007-4fb4-a17f-6c2dc8e53bbf · inbound

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data cites this paper.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.608391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:0f3422d97ce39e354ec930caefe4ab18733c918f312eb3bc07b35a6e8dde4a63

Observation 2383a679-d435-4953-9f81-f7ece23e5234 · inbound

Self-Distilled Policy Gradient cites this paper.

Self-Distilled Policy Gradient SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.906762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T11:24:11.878439Z digest=sha256:bd3d27f865d937056e38509cf25d3f40c5974611c09b6112733788bf0f3c112d

Observation d12d0f00-9328-4da1-be46-364c0af13427 · inbound

Reinforcement Learning from Rich Feedback with Distributional DAgger cites this paper.

Reinforcement Learning from Rich Feedback with Distributional DAgger SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:45.915273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T06:44:56.667364Z digest=sha256:14b76f1820625c065c1ab7640bcd953e511ccc2a60fe65b4559abd991c95bf34

Observation d8439ce1-be73-4c88-a3ba-b2bbd9d8e9e5 · inbound

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models cites this paper.

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.585403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T18:31:21.493677Z digest=sha256:ed36bf84bc47b85f1e0f6096492dc397b2a86e7984d49cae3fd055e30e718730

Observation 40290d98-3b46-434d-8d96-eb9c56da9001 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.902366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:2f17e949718020a9f097fd0129cf6b261390569220ac22958d8ecc7ac8ab30a8

Observation eaa9e059-daff-422f-8897-f28c490542eb · inbound

Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System cites this paper.

Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.837064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T21:43:01.949350Z digest=sha256:bc819973659610ec079020be33839b981ae4f20b6047aaab7a119a3321026d16

Observation 14275d2a-83f5-485b-b2f0-ea9c0e1c32d6 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.560962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:76b2a35059fce383b3dd1b2878e1b6ba017b2a34e339cf3a6b94a28b7d536f82

Observation 8829c247-63e7-43ce-b82f-14c140911edd · inbound

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning cites this paper.

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.429607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T04:49:28.598430Z digest=sha256:cf7eab1b220ab514a11e157384e8b9d8fff06108256f08c96e9bd89429a9ddcd

Observation 00ded8ff-7d33-420b-a0fd-003270913020 · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.387378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:e95f604314526b27a2ee809cc246dadfe5e72d9928208e2797c5c8399f1c3050

Observation 7c8263bc-0673-4a8f-83b1-d2cf75e97c99 · inbound

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs cites this paper.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:42.043252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T05:22:38.232552Z digest=sha256:42f6dd613058479e1a03142a3170e079fc4c2519b296ed55b725930c48ec6f71