Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:53:15.153126Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 18 inbound Pith citation observations for arXiv:2507.07017.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:53:15.153126Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.768850Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:09:43.539228Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9467682a-6022-4bbf-91ce-59a20c95f063 · outbound
First Return, Entropy-Eliciting Explore Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f96a7a7-cda7-4f5d-b81e-71a4897833e8 · outbound
First Return, Entropy-Eliciting Explore Using confidence bounds for exploitation-exploration trade-offs.JournalofMachineLearningResearch, 3(Nov):397–422, 2002
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation da3b65e4-0336-4c60-adc0-081ee30abc4b · outbound
First Return, Entropy-Eliciting Explore Finite-time analysis of the multiarmed bandit problem
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a73968-6880-4e1f-955d-dbbf5321f161 · outbound
First Return, Entropy-Eliciting Explore Language models are few-shot learners
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d58235a4-d96e-4d0e-888e-9f2403f8ca6c · outbound
First Return, Entropy-Eliciting Explore Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fa350c-2be6-4f5b-9662-3d9b8c654482 · outbound
First Return, Entropy-Eliciting Explore Process Reinforcement through Implicit Rewards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8df3f110-f943-4012-b73c-4441fdc1fff3 · outbound
First Return, Entropy-Eliciting Explore Assessing Validity of ICD-10 Administrative Data in Coding Comorbidities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0bfd072f-e7fb-4c9a-b297-695dacd32624 · outbound
First Return, Entropy-Eliciting Explore S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2a8d39d-42a5-4abe-bf7d-575a39e8159f · outbound
First Return, Entropy-Eliciting Explore Go-Explore: a New Approach for Hard-Exploration Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd151a39-067f-4f20-bc7f-7a61b8292e5d · outbound
First Return, Entropy-Eliciting Explore Stanley, and Jeff Clune
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2697bc8-ad10-49f3-ac3c-cd673bc6c02c · outbound
First Return, Entropy-Eliciting Explore A Survey on Mathematical Reasoning and Optimization with Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b4d0b61-8c05-4d5c-b784-da600d060f2b · outbound
First Return, Entropy-Eliciting Explore DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031a95bc-c452-4015-86f6-bdf98fce8502 · outbound
First Return, Entropy-Eliciting Explore Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fcc9aa5-d6d3-420c-90c1-532a17c17fe1 · outbound
First Return, Entropy-Eliciting Explore VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d42913-2226-4634-b56e-05c1653ec6c8 · outbound
First Return, Entropy-Eliciting Explore Temporal Sampling for Forgotten Reasoning in LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12059643-acbf-418a-8c11-f1bc77026d36 · outbound
First Return, Entropy-Eliciting Explore Let’s verify step by step, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c2f940-5cb4-4352-8f6d-a1c4c6f0d3b6 · outbound
First Return, Entropy-Eliciting Explore Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf567e20-b6ab-4828-be67-e36d925e32ff · outbound
First Return, Entropy-Eliciting Explore Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32634212-fcd9-4925-92e5-2cda5e921854 · outbound
First Return, Entropy-Eliciting Explore Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b41610-eaed-4414-b7d1-3c3323f7478a · outbound
First Return, Entropy-Eliciting Explore MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70828079-b558-4591-b755-4a9b04af1bef · outbound
First Return, Entropy-Eliciting Explore Training language models to follow instructions with human feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3d113656-ccc7-4533-8eef-555599ea29de · outbound
First Return, Entropy-Eliciting Explore A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 472d729b-e22a-4390-ae03-d9cb47d9f45b · outbound
First Return, Entropy-Eliciting Explore Qwen2.5 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d0e84f-522b-42c4-a44d-5483af15b854 · outbound
First Return, Entropy-Eliciting Explore Sequence level training with recurrent neural networks, 2016
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04ff86ca-7e0c-4c22-8e62-472b22cd7c6a · outbound
First Return, Entropy-Eliciting Explore Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be6ceca-94e6-444c-b349-47ec76ac118e · outbound
First Return, Entropy-Eliciting Explore High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a9bdeb-ac58-431b-a770-16c45ba907d3 · outbound
First Return, Entropy-Eliciting Explore Rewarding progress: Scaling automated process verifiers for llm reasoning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12d153fa-9aaa-46a8-8017-aa1bf075aa24 · outbound
First Return, Entropy-Eliciting Explore Spurious rewards: Rethinking training signals in rlvr
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b407463e-99a5-4f64-b345-a798cb0a5cf1 · outbound
First Return, Entropy-Eliciting Explore DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1628ed8d-f1db-4d53-8bdc-f2b7e54f2dc4 · outbound
First Return, Entropy-Eliciting Explore Hybridflow: A flexible and efficient rlhf framework
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20959a98-1f1f-48db-88a7-50c6fca993bd · outbound
First Return, Entropy-Eliciting Explore MIT press Cambridge, 1998
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f561fe24-6207-4082-83ee-b7881fc157b1 · outbound
First Return, Entropy-Eliciting Explore Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1–2):181–211, 1999
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation adba242a-f3f7-4a2a-86bf-1beaef2e5d1e · outbound
First Return, Entropy-Eliciting Explore Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d1700c0-2ca1-4f5b-9e34-db0a7d90a11c · outbound
First Return, Entropy-Eliciting Explore Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e43dfba1-b086-4b45-ab7b-4d162700aceb · outbound
First Return, Entropy-Eliciting Explore A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1537cf48-c20c-4d91-8a72-da3bbe84d475 · outbound
First Return, Entropy-Eliciting Explore DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bcd14f6-b579-4a79-ae0e-9ebfb33ca07c · outbound
First Return, Entropy-Eliciting Explore VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8755a62-5c3b-40d9-a1db-1e0eec0da0c5 · outbound
First Return, Entropy-Eliciting Explore SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588b069e-cf15-4971-b6ae-7fdd8a2ffc51 · outbound
First Return, Entropy-Eliciting Explore SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4534fe21-53e4-4202-91e8-dbccaba797ea · outbound
First Return, Entropy-Eliciting Explore Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a66a00-6f23-4a3f-ad4f-454fd58ab2dd · outbound
First Return, Entropy-Eliciting Explore Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b217b6-13f5-40da-9b5e-1d3ab295520a · outbound
First Return, Entropy-Eliciting Explore Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8e9ecb-8f18-4653-b893-fc7ccde05aad · inbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First Return, Entropy-Eliciting Explore
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b2af2a-9784-420d-848f-0382d14a4838 · inbound
Outcome-based Exploration for LLM Reasoning First Return, Entropy-Eliciting Explore
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31883257-a409-49e5-8385-e1ea6eae40f8 · inbound
Self-Reflective Generation at Test Time First Return, Entropy-Eliciting Explore
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e0d01b1-8879-442b-b095-39dc436c58a7 · inbound
Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts First Return, Entropy-Eliciting Explore
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa1a40d-e8fb-4d0d-a1a4-f25b78d4160c · inbound
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits First Return, Entropy-Eliciting Explore
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f373a8cc-ce6f-4c16-8b6d-3b79152a151b · inbound
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors First Return, Entropy-Eliciting Explore
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f689fb8-57bb-4870-a57f-205763537f54 · inbound
Selective Off-Policy Reference Tuning with Plan Guidance First Return, Entropy-Eliciting Explore
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6617a73d-f319-48fa-873a-3f4d3d1feda7 · inbound
Selective Off-Policy Reference Tuning with Plan Guidance First Return, Entropy-Eliciting Explore
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6962d0cf-3be1-46a4-9016-cd6295a98b42 · inbound
Reasoning Can Be Restored by Correcting a Few Decision Tokens First Return, Entropy-Eliciting Explore
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9099c6b-26ae-4bd5-b61a-9248ff6573f5 · inbound
Trust Region On-Policy Distillation First Return, Entropy-Eliciting Explore
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 024e228f-8c3a-40b3-93a7-ae7df8bbde22 · inbound
On Advantage Estimates for Max@K Policy Gradients First Return, Entropy-Eliciting Explore
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5523e626-d46b-4a60-8e80-73862cbeafab · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation First Return, Entropy-Eliciting Explore
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1254643d-d89f-4eb6-9f53-d3e6c4f518c3 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning First Return, Entropy-Eliciting Explore
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c977e00-2de3-42ac-b623-1fe22de75659 · inbound
Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier First Return, Entropy-Eliciting Explore
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b505331e-ed58-46a2-ba6d-ce078e8d855a · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning First Return, Entropy-Eliciting Explore
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 855d224b-6a44-4a97-ae37-fd7df83fc914 · inbound
What are Key Factors for Updates in RL for LLM Reasoning? First Return, Entropy-Eliciting Explore
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8331d4b8-dc11-4eda-b24c-c560ab2b1b34 · inbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization First Return, Entropy-Eliciting Explore
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8114cd00-53cd-4efb-baa9-9557b6121c75 · inbound
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning First Return, Entropy-Eliciting Explore
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.