Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:13.165308Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 35 inbound Pith citation observations for arXiv:2505.12346.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:13.165308Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:05.170825Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:49:52.427933Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2ee99a11-40e6-45e4-bc87-f4d4f5fe37d7 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93197f6b-9e95-4cab-a620-cd82fcce7263 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806637a3-f613-4a80-90d1-8ea4896e31ef · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72c7687-aa36-4f31-9f98-c9dde95e9424 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 060360f2-1a43-4706-b4f0-d4dbcd2858d1 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2502b848-428b-4b18-bfe3-feaa41f62102 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f557445a-6657-4e71-8261-a5188b2cbcd7 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7825dc23-4a1e-4b4f-85c6-2153f653c563 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788c11cf-4a69-464b-a656-4f9e647dda6b · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c595027-37fa-465e-83eb-0a25c13f14b6 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f197b82-6299-4f5e-ba2f-a2d960eaec36 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e992d9d-dbcf-4ad6-ae0e-c713228e4c24 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d74b4e0c-42e0-49e2-a291-df256a281dc7 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gemini: A Family of Highly Capable Multimodal Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c306b2-9fc6-4406-96a2-3acc9fc7fdf8 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization The claude 3 model family: Opus, sonnet, haiku
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5677145c-b958-48a5-8a0f-6cfe2a052da7 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b33c8d-64ff-47c5-883b-a9402670cbb0 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9ac5c6f-f0bf-4be6-8bb9-114314a10633 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d41709b-58e5-4e78-9b78-03a7a7b7d1a2 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Grpo-lead: A difficulty-aware reinforcement learn- ing approach for concise mathematical reasoning in language models.arXiv preprint arXiv:2504.09696, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e58237ea-71b5-4bb8-b946-dbe383eb6477 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411aa082-0154-4bea-993d-b0c547b36a80 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da99666-9fcb-4d92-a774-359878e57ad3 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 62ca4571-c604-40a6-98ad-08bf2d5c3b01 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Griffiths, Yuan Cao, and Karthik R Narasimhan
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4de2cb60-2685-4eb1-94b8-529d495a03cc · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fe917928-6605-40af-9bea-2bc25baec380 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb3a471-1774-4bd8-968a-38b86e82858b · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Curriculum learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656e5340-d350-417b-b530-0b29a7b9794d · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A mathematical theory of communication.The Bell system technical journal, 27(3):379–423, 1948
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0597bbb4-e21c-4d68-a9f7-20826ecbbe40 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Chain-of-thought prompting elicits reasoning in large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b47959-eda1-4b51-91e6-1a27791f76c6 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization LIMO: Less is More for Reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80079864-e5b5-42e5-8828-cfe4073c139c · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization ReST- MCTS*: LLM self-training via process reward guided tree search
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be79bc08-9e40-47b3-818d-42df74bdf59f · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53cb17cb-c18c-4078-8e31-ddab9477267e · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cc60ea2-ddc0-474e-9acc-a62709c82019 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Measuring mathematical problem solving with the MATH dataset
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 80a5c625-2dd1-43d7-bb23-158363ec38c9 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Solving quantitative reasoning problems with language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c608ec6-defa-464d-9b97-74487315f1df · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent AI
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae2d74c8-0d7e-42d9-83b0-244a2a88dffc · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879da7f4-6451-4c9d-8d98-f5e0fac47948 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2527d3e-6789-464e-b0c6-c7f67fa7a817 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18e2a4c-53a7-4fbb-91f4-f5ad7d9d0b9e · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c90c4c-6483-41a1-bd50-b4ab1ce0c251 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Evaluating Large Language Models Trained on Code
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee56f8a5-890a-45f2-9dae-01ef05def261 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22926a8-31ba-4a2b-9709-4450195aa258 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69e40fbb-3f69-438d-b86e-5541bab89966 · outbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4aba80-d73a-4225-914e-cd2b8ce53e20 · inbound
Reinforcing Video Reasoning with Focused Thinking SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b9f6000-6e9b-48f9-87c5-38a7c22a23d0 · inbound
Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f79782-9ec8-4f33-83a3-77ede35621bd · inbound
Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b01b73-13c2-4459-890e-334ee896733b · inbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892fbaba-8018-4613-907c-bbe0cdaaa6cc · inbound
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 1988
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7da6ce-8da8-4a9a-8688-18fe4b8d55e8 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · inbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe99f8e-e2ae-4851-85b9-d275d025b811 · inbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d457360-c1b1-41e6-8da8-77d54a0f8373 · inbound
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f273ed6b-ec28-433b-b6e8-b982d2353ee8 · inbound
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b031acc-1e60-4a7e-9592-a5d988094944 · inbound
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da19f184-186d-4aed-b54c-74ac41b3cf70 · inbound
Self-Distilled RLVR SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97aa497d-3246-4818-8110-ce4dfb077a56 · inbound
LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation caa9ebf4-db47-47db-b183-b28db1c89bbc · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40b4c59d-42a6-46d5-a2f3-85242df039df · inbound
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 131bf09f-eb00-4823-b093-0ab1106f3131 · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c9ae2cc9-c05b-4afd-b031-5b719c166158 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ddcc6ed-9c0a-4a53-b513-ecf095e389f3 · inbound
Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c96d400-c4e2-4a8f-8d3c-f44cd106e081 · inbound
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7f265df0-44b6-4a12-a82c-653cfcab9e64 · inbound
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation da4eb23b-efe3-4d94-bfdc-eed8427c032c · inbound
Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94c188f2-641f-4ea3-8252-6faeb9c094f9 · inbound
Holder Policy Optimisation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0a337026-3afb-47eb-956a-f88f1dcfd852 · inbound
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82fbf2f1-0959-46e2-a70b-f94cd95608b1 · inbound
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45242034-0215-4725-8b8e-0e4ae85a8d30 · inbound
Trust Region On-Policy Distillation SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7416aa03-8007-4fb4-a17f-6c2dc8e53bbf · inbound
FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2383a679-d435-4953-9f81-f7ece23e5234 · inbound
Self-Distilled Policy Gradient SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d12d0f00-9328-4da1-be46-364c0af13427 · inbound
Reinforcement Learning from Rich Feedback with Distributional DAgger SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8439ce1-be73-4c88-a3ba-b2bbd9d8e9e5 · inbound
Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40290d98-3b46-434d-8d96-eb9c56da9001 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eaa9e059-daff-422f-8897-f28c490542eb · inbound
Empowering Economic Simulation Through Situation-Aware Llm-Driven Generative System SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14275d2a-83f5-485b-b2f0-ea9c0e1c32d6 · inbound
What are Key Factors for Updates in RL for LLM Reasoning? SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8829c247-63e7-43ce-b82f-14c140911edd · inbound
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 00ded8ff-7d33-420b-a0fd-003270913020 · inbound
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c8263bc-0673-4a8f-83b1-d2cf75e97c99 · inbound
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.