Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:54.829237Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2509.00125.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:54.829237Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:33:43.695859Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:40.662689Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0fae64bf-bc87-4977-87d0-02d7f71eaf7e · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Learning to reason with llms, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9202f34b-3cb7-442f-97a4-4eab7dba5831 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd8135b-080c-4eec-bf07-a9e4ab3cbfe9 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c9fb5e-b0bb-4a94-9879-00c3ce7d357a · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Qwen2.5 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b20f8d-b312-4c59-b42d-59e511803bf2 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Let’s verify step by step
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb93fe45-4a55-4b3c-8e64-3b59a4fb0a79 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3703a958-d47e-45d0-92bb-5518f6b39972 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Scalable best-of-n selection for large language models via self-certainty
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21250019-abda-44ec-8a32-2a4851ebe388 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The impact of intrinsic rewards on exploration in Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea1186f2-c4dc-4e56-a502-2c850b501999 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Thompson sampling: An asymptotically optimal finite-time analysis
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71f55c8d-7cfa-4b66-b4e4-a509853d5451 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Thompson sampling for contextual bandits with linear payoffs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4060531c-1847-4585-b001-198f32c05dd6 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Exploration-Exploitation in Constrained MDPs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822ae0fc-6db9-407b-be99-fa09a6948950 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Why is posterior sampling better than optimism for reinforcement learning? In ICML, 2017
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1e214fe-a27f-4f1e-8e68-16d1520ea192 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First return, then explore
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5cfca35f-e4c9-4323-aab3-7ff3f6200c07 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Automatic goal generation for reinforcement learning agents
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0896c7a-5d24-49cc-99f5-bbf3454a21f9 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning MIT press Cambridge, 1998
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b22afa6-dbbf-4e5e-a976-d323cbf41029 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Adaptiveε-greedy exploration in reinforcement learning based on value differences
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation af537e40-7b37-47cf-8762-26f1ddde1bed · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning # exploration: A study of count-based exploration for deep reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 625027fc-5657-457b-bc6a-a25ba0a7f7ea · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Deep exploration via bootstrapped dqn
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47b9055b-4018-4590-9ea4-cd41539fb231 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curious model-building control systems
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f64fc493-eefc-44ad-931a-d5febad69903 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curiosity-driven exploration by self-supervised prediction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 071b3708-a6bb-4d24-9a98-dd9e66e6f994 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Exploration by Random Network Distillation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b194bb3f-0f88-4cef-83ff-3c494a524657 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Planning to explore via self-supervised world models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 954ec9c2-13a1-461f-8f5b-5848e91dbe33 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Vime: Variational information maximizing exploration
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7cc4687-085e-40d0-bed4-45c222baf6b1 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Diversity is all you need: Learning skills without a reward function
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d57a21b2-6589-4b9f-a40b-01630ae61e98 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Ttrl: Test-time reinforcement learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9420cb17-f685-4dd4-a6d8-48a6600f33fa · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Learning to Reason without External Rewards
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747bd84f-aaff-4c49-8042-15dc9311dd0b · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c343ed-07cd-4656-85ed-de4239c3f1c1 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72352bc2-0201-46b3-9b37-71dd8e686483 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning One-shot Entropy Minimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2e1034f-eb4b-4ef7-9838-69608525c0c4 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8e9ecb-8f18-4653-b893-fc7ccde05aad · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First Return, Entropy-Eliciting Explore
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f036f048-214b-4d9e-8a89-9046536ab732 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Reasoning with Exploration: An Entropy Perspective
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77947edf-38ea-4181-93e3-f094562744d5 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c944a4df-f17b-4859-8da6-eafd9de83834 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Free Process Rewards without Process Labels
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ace7843-b82d-4429-a83c-473e3d336d9d · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3a05f2-725d-472a-89fd-6cf0e83dd2eb · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bcc21ec-a01b-4b73-b936-4c3833198956 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a84f266-0bad-4488-ae18-149dd2857216 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7d0dc9-5dd4-47be-8d47-40168e4891c3 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curriculum learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab1850b-1883-4ced-b777-0144ee09d814 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Efficient memory management for large language model serving with pagedattention
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3480521a-824d-43ae-a716-80da712b82e7 · outbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Math-Verify: Math Verification Library.https://github.com/huggingface/math-verify, 2025
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ccf703d3-489a-4cbe-adaf-5d000ca5b6bf · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09b23c32-fb30-4c1a-a096-9dbebe7a4b2b · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c05e544-ae55-45ef-b807-6a1da0229726 · inbound
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4f5a4c4-890c-4235-bb9b-38ae12da4741 · inbound
Epistemic Uncertainty for Test-Time Discovery Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation efaff5c3-2482-45f9-8ded-02939fa1f10b · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09d9d026-81c7-47ef-bb80-983a28743583 · inbound
Formalizing Task-Space Complexity for Zero-Shot Generalization Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7648875b-3f7a-4c46-af90-34ae6c91614b · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.