Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:59:14.608524Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 25 inbound Pith citation observations for arXiv:2509.06941.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:59:14.608524Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:16.022969Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
38 of 38 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 772784dd-2c91-4acc-97f1-975b46d09b0d · outbound
Outcome-based Exploration for LLM Reasoning 2:fort= 1,2,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 288c2f09-c7eb-49ea-adb0-38ff0434e55b · outbound
Outcome-based Exploration for LLM Reasoning Dataset Reset Policy Optimization for RLHF
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36f566e-f01d-46dc-9179-90ccd70433a4 · outbound
Outcome-based Exploration for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ba54e2-b4ab-4ee0-846f-2aa771c9ce62 · outbound
Outcome-based Exploration for LLM Reasoning Reasoning with Exploration: An Entropy Perspective
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce41058-e781-4281-a027-09c512726caa · outbound
Outcome-based Exploration for LLM Reasoning Weight ensembling improves reasoning in language models.arXiv preprint arXiv:2504.10478,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0d72bb-28eb-42d3-b440-06304a84dbfb · outbound
Outcome-based Exploration for LLM Reasoning The Statistical Complexity of Interactive Decision Making
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6109ad-08bf-4aef-a875-997ebe7deb32 · outbound
Outcome-based Exploration for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a506f41a-267b-49c8-a223-5e200306f839 · outbound
Outcome-based Exploration for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359e9c1c-7d9d-440f-96fe-bbdbc4be1d7a · outbound
Outcome-based Exploration for LLM Reasoning OpenAI o1 System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94662f91-9583-418f-89b4-06d067697620 · outbound
Outcome-based Exploration for LLM Reasoning Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c1a1a68-f682-4f72-ab3e-c7d1d3067f65 · outbound
Outcome-based Exploration for LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a22718-0a52-4f80-9d4b-b5cd772f1e00 · outbound
Outcome-based Exploration for LLM Reasoning Approximating kl divergence, 2020.URL http://joschu
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 498f03d7-ccc7-4751-8b18-066d62061239 · outbound
Outcome-based Exploration for LLM Reasoning Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c7fea0-23d2-4572-910d-e602d65a79a6 · outbound
Outcome-based Exploration for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2daa3bfb-6428-44ca-a5e3-92d2dcbcebee · outbound
Outcome-based Exploration for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae28b6aa-0078-4384-a72b-596394671fad · outbound
Outcome-based Exploration for LLM Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d87d62-3600-436e-8f4c-b412c9fef2d6 · outbound
Outcome-based Exploration for LLM Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb93e9e2-8fda-416f-af68-8a2d11cd188d · outbound
Outcome-based Exploration for LLM Reasoning On a few pitfalls in KL divergence gradient estimation for RL
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d78bad3-7a1d-4923-8247-095ae001384d · outbound
Outcome-based Exploration for LLM Reasoning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab716baf-2fff-4450-87df-83fb63bd4183 · outbound
Outcome-based Exploration for LLM Reasoning The invisible leash: Why rlvr may not escape its origin.arXiv preprint arXiv:2507.14843,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee64cf97-aa98-409a-94f1-1a11ad286fef · outbound
Outcome-based Exploration for LLM Reasoning Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca620d9e-6f09-4fc7-bfbe-5a69903cbd8f · outbound
Outcome-based Exploration for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 635fd568-f897-4a7b-80ec-6f02b0858d77 · outbound
Outcome-based Exploration for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac98f467-6180-45da-b120-1910ac200a53 · outbound
Outcome-based Exploration for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02725c34-f758-487b-923c-88331c46836e · outbound
Outcome-based Exploration for LLM Reasoning The Price of Format: Diversity Collapse in LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07e30d42-0401-4a3e-98d7-9f776f3d8e3e · outbound
Outcome-based Exploration for LLM Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b2af2a-9784-420d-848f-0382d14a4838 · outbound
Outcome-based Exploration for LLM Reasoning First Return, Entropy-Eliciting Explore
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923ed1ad-05bf-484b-b01e-fd396b1f5ffd · outbound
Outcome-based Exploration for LLM Reasoning Optimistically optimistic exploration for provably efficient infinite- horizon reinforcement and imitation learning.arXiv preprint arXiv:2502.13900,
Reference 1996
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b198d2-3e16-4493-851b-f3f8ea1ea061 · outbound
Outcome-based Exploration for LLM Reasoning Exploration by Random Network Distillation
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8590e1fa-3675-4041-bd44-43f2f9dd1d6e · outbound
Outcome-based Exploration for LLM Reasoning Diverse Preference Optimization
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c308a8-ab1d-4242-83f7-f1ad2fa216db · outbound
Outcome-based Exploration for LLM Reasoning Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc9057c-9b41-48ca-8c6c-47c9a99027ba · outbound
Outcome-based Exploration for LLM Reasoning Qwen2.5 Technical Report
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e79bd9-5de4-4b91-b579-44a002b56daf · outbound
Outcome-based Exploration for LLM Reasoning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4927a9db-db97-43b7-bc2c-0629bb70414b · outbound
Outcome-based Exploration for LLM Reasoning Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04a1e4f-1a34-405c-bc14-f9f4ca59e70d · outbound
Outcome-based Exploration for LLM Reasoning Attributing mode collapse in the fine-tuning of large language models
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 964e69e0-4b1d-4850-ae05-6aa530a6b73f · outbound
Outcome-based Exploration for LLM Reasoning Asymmetric rein- force for off-policy reinforcement learning: Balancing positive and negative rewards.arXiv preprint arXiv:2506.20520,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61d8773-626d-46fe-b3bf-f7fa28628346 · outbound
Outcome-based Exploration for LLM Reasoning Online Preference Alignment for Language Models via Count-based Exploration
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b385f625-764e-405a-b34c-728dee68553d · outbound
Outcome-based Exploration for LLM Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba17bf7-aa41-4982-be33-c74970665f6d · inbound
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Outcome-based Exploration for LLM Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93fc4f7e-c621-4a05-9be5-d73b16ef952b · inbound
Polychromic Objectives for Reinforcement Learning Outcome-based Exploration for LLM Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7bb58a8-660f-45d5-a3b7-88336b8b09c5 · inbound
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs Outcome-based Exploration for LLM Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47bb9eb1-86ab-484e-aecf-a73e478b05a3 · inbound
On the optimization dynamics of RLVR: Gradient gap and step size thresholds Outcome-based Exploration for LLM Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f7150d8a-8e2d-4f20-95f4-cc2b6b036feb · inbound
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective Outcome-based Exploration for LLM Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8365833d-3a81-4807-8ea1-032cec879060 · inbound
Beyond the Sampled Token: Preserving Candidate Support in RLVR Outcome-based Exploration for LLM Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f61e66d-647e-4ac4-9bb8-3c2b623fc48a · inbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec353b52-5182-4790-b97c-e9927c819038 · inbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bc8eb1c4-2005-4bcb-8c0a-d00995cf9619 · inbound
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Outcome-based Exploration for LLM Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3a0da2d4-b844-4ad7-966f-e3e171b50b9c · inbound
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Outcome-based Exploration for LLM Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9d1c9fdd-fa13-48a5-8ad1-5b2657b1e058 · inbound
Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models Outcome-based Exploration for LLM Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 15b83eb1-f291-473d-bba4-8dbfa2f28860 · inbound
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning Outcome-based Exploration for LLM Reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b0ad7f3-33f2-4b9e-a5a8-d3c0de4561b9 · inbound
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR Outcome-based Exploration for LLM Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c2f4873-9732-44ca-92cd-a2ca7728e82b · inbound
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Outcome-based Exploration for LLM Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cfbde808-e086-48c3-a262-22ea20a23115 · inbound
Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Outcome-based Exploration for LLM Reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 04a2bc2c-32e2-40d3-9bb4-51b9022ca6e9 · inbound
On Advantage Estimates for Max@K Policy Gradients Outcome-based Exploration for LLM Reasoning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8126bf8-e979-40d1-a9e8-ef3344887b16 · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Outcome-based Exploration for LLM Reasoning
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9b42875-8786-441d-b3f2-9035010c79b4 · inbound
Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling Outcome-based Exploration for LLM Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 460713b0-7beb-42ad-921e-862a533385bf · inbound
Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning Outcome-based Exploration for LLM Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09bd86c5-8aec-4a16-8f1c-288d66fe7ce6 · inbound
When are likely answers right? On Sequence Probability and Correctness in LLMs Outcome-based Exploration for LLM Reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9ae9caf7-aae5-4d94-965c-b29c28373d27 · inbound
Depth-Entropy Guided Sampling for Training-Free LLM Reasoning Outcome-based Exploration for LLM Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a903df79-6c71-478c-beec-f75b19897c0a · inbound
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches Outcome-based Exploration for LLM Reasoning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f9f5bd-daf5-4f16-aa10-d070be0d09a1 · inbound
Bridging Compute- and Data-Optimal Pretraining Outcome-based Exploration for LLM Reasoning
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ca818a-4402-4f00-a1c5-b184bffdb329 · inbound
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Outcome-based Exploration for LLM Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0751ac2f-e549-4bfc-917f-80d0aead3b5c · inbound
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Outcome-based Exploration for LLM Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.