Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:33.393958Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 11 inbound Pith citation observations for arXiv:2506.09026.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:33.393958Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:41.937333Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
92 of 92 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 563ffcf9-a85f-4873-935f-8451d9f9cbf1 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs On the theory of policy gradient methods: Optimality, approximation, and distribution shift.Journal of Machine Learning Research, 22(98):1–76, 2021
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef38348-1375-4b6c-84e8-f52d09fd0b6d · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dadfa8ab-d078-4d79-bdbd-7b99b5afe111 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75d7da3b-4d43-4d9b-8a91-a506884db8d7 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8abdacc-97b9-40be-8a4e-b4f4c1bcec04 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bfbdea-a97e-407a-aa9e-0084dca6ec06 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50086ad2-4bd5-4074-a2df-9092189d01b8 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dbd99e9-5309-4290-99df-a3a249b670c8 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Stream of Search (SoS): Learning to Search in Language
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b874ac4e-68f7-4292-9003-d210c3136e34 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc605c3a-310f-4ff3-9b8f-329b499d21c4 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5601a49a-b4a1-448d-b99c-c8f541ae153c · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee39a6f-de9c-45e7-b557-533f155b25a5 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unsupervised Meta-Learning for Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be508c46-43de-44ef-ad57-19332c637362 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Provably Efficient Maximum Entropy Exploration
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa5bae73-d49d-4d36-a620-4458f701bdf4 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2244849d-c21a-4ff2-82fb-0ec8ec7e670a · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Improvement in Language Models: The Sharpening Mechanism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ff2b50-c5fd-4950-a02b-65d83d0fd24a · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Evaluation-time Compute with Reasoning Models as Evaluators
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14075a3-af6b-4cea-a2df-df48b7991ca5 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Can large language models explore in-context?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a9a868-d0cd-4408-b1a8-34f7e025060a · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Training Language Models to Self-Correct via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6951c05a-5e10-4317-a2a4-0976c4bd09f1 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Understanding the Complexity Gains of Single-Task RL with a Curriculum
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0473c9c2-a41b-4010-bf25-5fd125a256b7 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Long-context LLMs Struggle with Long In-context Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d64ab7-2198-4b1c-a6c5-4b0a9ea16206 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Learning Abstract Models for Strategic Exploration and Fast Reward Transfer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfe38482-8f06-465a-97f4-0d92e1ae9e09 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327cfcbf-ed19-48cf-9796-ee89fe80cad3 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Understanding R1-Zero-Like Training: A Critical Perspective
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b53ecd-4144-419a-8716-247b70a4f766 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Acemath: Advancing frontier math reasoning with post-training and reward modeling.arXiv preprint, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92b2fbf2-884b-45a3-b961-3db8c7a8d3c7 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Deepcoder: A fully open-source 14b coder at o3-mini level, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4aa8473-3a95-48d6-8c14-8069af19c0fa · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10aa9ec4-ba0e-43c5-90e6-19807c442f98 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs s1: Simple test-time scaling,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52911944-a89a-47e7-ae74-3accc39b8d79 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22a6a3f-f6ff-4d52-b80b-d5f983955527 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs s1: Simple test-time scaling
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ee57e0-46a5-4534-b3f4-40b377ee033f · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Maximizing Confidence Alone Improves Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7788e3dc-c34d-47d8-9433-c4febbdd1c91 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs OpenAI o1 System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9501d6f4-81d4-425f-9b26-ef0835d084cd · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing anytime reasoning via budget relative policy optimization.arXiv preprint arXiv:2505.13438, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82baf72a-03f8-433d-926d-851f97505323 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs John Wiley & Sons, Inc., 1994
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83245270-4548-4a11-8767-ddf00921963f · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93881a94-72c1-4b83-815f-b6aa2f1a469a · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1782b1e6-2e5d-4e88-bd43-c2cc14ef8572 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7082dd8b-ab49-4565-b43e-4b6fa9fd41ce · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d155d0d9-b405-46c5-9231-8f57d771e0eb · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Test-Time Compute Without Verification or RL is Suboptimal
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba954c4-8f3f-4ec8-8b34-bcad6404e7b1 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Opti- mizing llm test-time compute involves solving a meta-rl problem.https://blog.ml.cmu.edu/,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93354397-44c3-457a-8c2b-ad6a5463736c · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Spurious rewards: Rethinking training signals in rlvr, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed5a159-d96f-420e-8049-51292198c6ff · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Can large reasoning models self-train?, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28761a3-3781-4ea6-ba56-de81838ac795 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs HybridFlow: A Flexible and Efficient RLHF Framework
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72ebdd1-a06d-45cf-ae99-5f5e48a5f708 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3773da43-37af-48b7-9350-02da630331d1 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 207ae8f2-58cf-4814-a505-d16add4c7125 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d722aab7-75a3-4e2c-b6c7-a7e5fb01397f · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b760d7ee-a2d3-4a3e-a999-b4747dad6391 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41efc73e-f3a3-4f0f-a89c-97b4007aa361 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014f4051-8c67-462c-929d-4f6ef5cb1548 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs All roads lead to likelihood: The value of reinforcement learning in fine-tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741f74ff-94a1-47c6-a190-cdf584061006 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Open Thoughts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878ef22e-f7fb-4336-9912-49ecb9b789eb · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f835cd85-0c80-4003-b734-4115945c1f75 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ac5cef-66ee-4c4d-9405-8c67ec8399ff · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0eaa6fe-7121-44d7-a84b-6ce3418d5c79 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1e103d-0c2b-46f9-9ec7-68d171daa5b5 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9f2196-f3c0-4924-9bb5-06e05a112b7d · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968c06be-1c7f-4e33-a8d9-a2e465a1f2d2 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Qwen3 Technical Report
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b61a26d-890a-4211-83ab-ee5887694417 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe2ba65-17b0-4fc9-a82c-421acd9a1363 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26d30ab5-fa15-4fc6-bc30-355c07bb69e5 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233747cc-2482-4e80-a50b-534029d8f11f · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7f6153-5475-49de-8a4c-b66f7025612a · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Star: Bootstrapping reasoning with reasoning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7413e4-ec3c-41e5-866f-878829bce260 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Learning to Reason without External Rewards
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5d42ea-cccc-42f2-9151-fa579a2a658d · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b24c0e-a7a5-4ddb-bdfa-0e4246605f7a · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Looking at the other numbers: 77 - 70 = 7 97 - 73 = 24 (interesting, we already have
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 73793d47-d65d-4c6b-be67-7ad5aa220217 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16ec9fdc-ddb2-4d41-8eec-2bc04cbf672b · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs We need to get from 37 to 466, which means we need to multiply by 12.5
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cba0c388-9ba8-48ee-91d6-0e00a7c42aaa · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7703ae3e-ae22-45dc-a9ec-624e83181129 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38c15b0c-bec0-4193-848a-65acd990faf9 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b8878db-da0e-463d-b3f2-5872b395b83c · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48455646-a978-45c3-8a4c-22f364a55508 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bcb27548-26d4-4336-8b9f-2c7770a86ee7 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a42f1119-c124-4ced-840f-64137e979db0 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66d47b82-d388-4183-a294-7b9d7472290d · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 476391ed-9599-46d5-ac75-e69fa5892a14 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f1d463a-e987-489c-99e0-3e57a75edd6e · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 75ba2295-f46d-4876-a35b-e71b4b53bfff · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66dfd81e-85ea-4ea7-bdde-1fa303f6a108 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2e2a2a8-3ad0-47ab-9d10-8fe797d320e0 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 884f3572-8433-4f42-9e4b-22f75dd0c333 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 437b83b2-677a-4490-8b68-89b6d5c1cd3c · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 269eedb9-3091-40f7-8272-7aefe649550e · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c5a907d-1706-4adc-88f4-31f4dc9dbfbd · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs 31.5 (not helpful)
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7db092d6-5cdf-4905-bf33-b562e09ffb1b · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd6f2c5d-3087-473a-85c4-ca16269bc216 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8fa66cf5-8f1c-40b3-b0a2-9e5cf583b650 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2efbcaa7-2bed-4693-941a-5eb70b13dbd1 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 415a2322-a148-4db3-849e-c71a98eef83b · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Hmm, let me think about how to approach this
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9759744a-aa27-460a-9737-6bad422b14a8 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e8a353ca-341f-4290-925d-faf4647054df · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Again, multiply 347 by 5 and add two zeros
Reference 500
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ca58a45-cbde-464b-9c44-32b754e36ef5 · outbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9749439b-2f45-4577-a302-9162374af12e · inbound
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc96f34-873c-4b75-ae35-482082edbde9 · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e79bd9-5de4-4b91-b579-44a002b56daf · inbound
Outcome-based Exploration for LLM Reasoning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1e0385-5950-46db-a390-5826abdc1213 · inbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac25798b-d911-4be4-9de9-432a70e14263 · inbound
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 96a65a17-114f-4a3b-9585-6ae340cd7f01 · inbound
What Does Flow Matching Bring To TD Learning? e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97b7fe68-7dfb-4165-bc4e-0cade0795304 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11b5d117-b3eb-4c77-8be0-13e8c12225c7 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9de63b67-d0db-426d-ae15-87987fce5a98 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 978bf51c-3054-476a-bfff-6e3746dd43dd · inbound
On Advantage Estimates for Max@K Policy Gradients e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 18fc5695-e22c-4d53-9796-fff3c94285ae · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.