Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:03:28.188682Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2508.07534.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:03:28.188682Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T00:02:25.252362Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 322ba5f0-cb9c-49f0-8b58-317bba6f33f2 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2905d31a-65ad-473e-818e-f86ae11c2154 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR A Survey of Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e3e04c-cc53-4260-ba52-1bdd66cd14ea · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Training language models to follow instructions with human feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b421bde8-3511-4eac-92c2-d25e8ddbe07a · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9624f8e9-1b9d-4e39-bf3b-54a9be2b8fbc · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4095a71-a846-4d77-bd8c-352760ee80b8 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24bdff8e-b204-4b53-a0f5-c33bd7634669 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Reasoning with Exploration: An Entropy Perspective
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011fa54c-4c8a-4470-a3b9-79a31d60fabc · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Exploration in deep reinforcement learning: A survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4e661c4-d523-4537-b298-07d6ce18cea9 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67bce6db-ecda-40bc-9b2e-5cd71a80e58a · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d691cb-a13d-487d-b724-d22111c95a82 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The surprising effectiveness of negative reinforcement in llm reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd046845-72d6-4737-8501-5894b9ebd3da · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2f1dcbc-69c2-4cb8-bdc3-a75f9d9db686 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50c3b81c-2c61-4c63-8326-dde21d646ebe · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b894881-4190-409d-9d68-3950e4797a83 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ca83fb-bc87-4b29-8023-c78c84d38c92 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Evaluating Large Language Models Trained on Code
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40f1373-96d7-47d6-bf88-a85fa7fb7c7d · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55958e60-b658-4583-af0f-e5401c34048b · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a748fea4-02d5-468e-bc1d-8767a55ab860 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Let's verify step by step
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54180ce1-e458-43c7-9ce9-475e1ca386df · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Generative verifiers: Reward modeling as next-token prediction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 162588cb-f6b0-4e4e-8c70-8508c9f4bc9b · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The lessons of developing process reward models in mathematical reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea005420-8fac-45fb-be45-2bb5acf2547f · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Enhancing LLM Reasoning with Reward-guided Tree Search
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e5d277-296f-40f4-95ee-93a817267c03 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Inference-time scaling for generalist reward modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b52097-3d64-4337-b9f5-4207d2bca587 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Heimdall: test-time scaling on the generative verification
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d8a914-54b5-435e-9846-bef7c95d6450 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Towards Effective Code-Integrated Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4915f496-4c8b-48cd-a4fb-93f90b371a93 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8f702a4-4f9d-4541-ab53-b33821fc020f · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab8723e-19b3-466a-9f43-335a577da71b · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Not everything is all you need: Toward low-redundant optimization for large language model alignment
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14e94b11-8757-47ad-a607-6a1e15eaeee5 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Towards a Human-like Open-Domain Chatbot
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68da4352-c443-4b1c-972e-2e4d88fc095d · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Maximizing Confidence Alone Improves Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f1999f-6b2b-455f-b6bb-c20fc753d306 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Proximal Policy Optimization Algorithms
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017bbefa-64cf-4014-b67b-1bb109992938 · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca777cb-58e4-41ec-8d83-733f6ba0eeca · outbound
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Agentic Reinforced Policy Optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a761aa7-9967-4ef2-81d7-1f3f78de71b1 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffad19ef-e6cf-4214-9ca7-868e40a51bff · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c712013-bb5f-41fa-9ad7-1ab6b67256b3 · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.