Pith. sign in

Paper Citation Record · LEDGER

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

As of 19 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 12 inbound Pith citation observations for arXiv:2506.09026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09026 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:33.393958Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:32:13.013301Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 563ffcf9-a85f-4873-935f-8451d9f9cbf1 · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.Journal of Machine Learning Research, 22(98):1–76, 2021.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs On the theory of policy gradient methods: Optimality, approximation, and distribution shift.Journal of Machine Learning Research, 22(98):1–76, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.231877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.231877Z digest=sha256:038187f3cda40a7e01a53b93ff7dcc86b09705c5d0c1d50e454eb96b2ddc44e8

Observation bef38348-1375-4b6c-84e8-f52d09fd0b6d · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.293292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.293292Z digest=sha256:83d1dd6827b70909cf7e03f514905aeb2acd097589e5888c11159832b5395cae

Observation dadfa8ab-d078-4d79-bdbd-7b99b5afe111 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.334902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.334902Z digest=sha256:2cb4447efa9f21fd1572b9938e005cfa62dce40895f2f04d963604dab116361b

Observation 75d7da3b-4d43-4d9b-8a91-a506884db8d7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.373350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.373350Z digest=sha256:0afeda1c57b137be78e19b1a40f3b54ee3cde39077ecefb38f80c5bbf807855a

Observation d8abdacc-97b9-40be-8a4e-b4f4c1bcec04 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.417829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.417829Z digest=sha256:211d64431b4b20b719ec0b98dd9f7b22f6aa970305be9ff4bda41842595c1e05

Observation 34bfbdea-a97e-407a-aa9e-0084dca6ec06 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.549636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.549636Z digest=sha256:b05e496baf735fe5ce1ad9b181bda6a73e522502980ddddcbf72bfe7a7d361ca

Observation 50086ad2-4bd5-4074-a2df-9092189d01b8 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.613833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.613833Z digest=sha256:700f6cdcc80fe486a3240493114e3b2aa65fcb21d0722b8d64b704abbe0a07b4

Observation 3dbd99e9-5309-4290-99df-a3a249b670c8 · outbound

This paper cites Stream of Search (SoS): Learning to Search in Language.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Stream of Search (SoS): Learning to Search in Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.660715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.660715Z digest=sha256:d1f17f3dbbb55ed60a5414fff74eda24471c88a4680236dfe4e8a58d8564cd00

Observation b874ac4e-68f7-4292-9003-d210c3136e34 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.755143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.755143Z digest=sha256:c1245ffbf831504e65f392d677dae4792ae04a7a34527bc958220afe2ac5c1ed

Observation dc605c3a-310f-4ff3-9b8f-329b499d21c4 · outbound

This paper cites Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.809592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.809592Z digest=sha256:4d9528bb3e606b7e3986c5a5cd362b9cbd974e66fe44c83b973a9acd6c22ecd5

Observation 5601a49a-b4a1-448d-b99c-c8f541ae153c · outbound

This paper cites Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.861972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.861972Z digest=sha256:bff0c6dd0d65f8b9c0276d633a7ba0ef1610338afcc45529fbccbf5307dd5672

Observation 5ee39a6f-de9c-45e7-b557-533f155b25a5 · outbound

This paper cites Unsupervised Meta-Learning for Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unsupervised Meta-Learning for Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.925948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.925948Z digest=sha256:fba1a83046819c4f03a075957deacbdf28d3c2b15093df174bbbd55c06fe41da

Observation be508c46-43de-44ef-ad57-19332c637362 · outbound

This paper cites Provably Efficient Maximum Entropy Exploration.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Provably Efficient Maximum Entropy Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.984106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.984106Z digest=sha256:4b233ed3acb208d7d63ead1e8de00caf6e30da7bd2a73c38c1373f05cedf7554

Observation fa5bae73-d49d-4d36-a620-4458f701bdf4 · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.061189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.061189Z digest=sha256:ae9f8523c275133dfd1428c3c07c67a2ef5ab78887dfb933bb533226664fbac3

Observation 2244849d-c21a-4ff2-82fb-0ec8ec7e670a · outbound

This paper cites Self-Improvement in Language Models: The Sharpening Mechanism.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.163556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.163556Z digest=sha256:80c1157b9543bae339920f07f5013c1a10fbef3238570f3f115a90a67792250e

Observation 55ff2b50-c5fd-4950-a02b-65d83d0fd24a · outbound

This paper cites Scaling Evaluation-time Compute with Reasoning Models as Evaluators.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.243355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.243355Z digest=sha256:9125804386306e3be4601086807a1de218f02869b58edb2e15404feb8fbc700b

Observation c14075a3-af6b-4cea-a2df-df48b7991ca5 · outbound

This paper cites Can large language models explore in-context?.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Can large language models explore in-context?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.307343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.307343Z digest=sha256:e195ad93d84732366cc794d0c5fec124515d0f0e8d40ed546eda74eef950c9c1

Observation 12a9a868-d0cd-4408-b1a8-34f7e025060a · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Training Language Models to Self-Correct via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.372345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.372345Z digest=sha256:f86a1db1b6f257c8c216218bb0e6ea254edf3b6fe495e3aff87aef83e97a1940

Observation 6951c05a-5e10-4317-a2a4-0976c4bd09f1 · outbound

This paper cites Understanding the Complexity Gains of Single-Task RL with a Curriculum.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Understanding the Complexity Gains of Single-Task RL with a Curriculum

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:03:34.373383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:26.430721Z digest=sha256:740e612eb5086fdf4eca47068e8c996f76b6187a292e492148170073387e11aa

Observation 0473c9c2-a41b-4010-bf25-5fd125a256b7 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Long-context LLMs Struggle with Long In-context Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.464454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.464454Z digest=sha256:91be2b5fce1ffd5a98c55f15572ed6ea72ff8f12be9788fee8e3f035725a3c3a

Observation b2d64ab7-2198-4b1c-a6c5-4b0a9ea16206 · outbound

This paper cites Learning Abstract Models for Strategic Exploration and Fast Reward Transfer.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Learning Abstract Models for Strategic Exploration and Fast Reward Transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:03:34.135617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:26.505425Z digest=sha256:5bd51531ffa241073d794decf1eb1451e439adbccf2cca0ed83ac72ef3d3521f

Observation cfe38482-8f06-465a-97f4-0d92e1ae9e09 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.538629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.538629Z digest=sha256:8333a2d4ecd4ac9055a8b48af86521ed0d569711058ee9643d3c7e17833265c1

Observation 327cfcbf-ed19-48cf-9796-ee89fe80cad3 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Understanding R1-Zero-Like Training: A Critical Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.569626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.569626Z digest=sha256:dd880a62e03bb2753d338766d50a58440d74aee330164041ca2a51cba17c5a9f

Observation 09b53ecd-4144-419a-8716-247b70a4f766 · outbound

This paper cites Acemath: Advancing frontier math reasoning with post-training and reward modeling.arXiv preprint, 2024.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Acemath: Advancing frontier math reasoning with post-training and reward modeling.arXiv preprint, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.638436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.638436Z digest=sha256:5a8012202a7e3d58915925db92ea8443ec89de3d1a442b39a0db4103a09be0b0

Observation 92b2fbf2-884b-45a3-b961-3db8c7a8d3c7 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Deepcoder: A fully open-source 14b coder at o3-mini level, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.690066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.690066Z digest=sha256:4d33a194d73463aba90e71225c14a2bd960126f6297fef6e5d522ebbc74b2ce0

Observation e4aa8473-3a95-48d6-8c14-8069af19c0fa · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.747061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.747061Z digest=sha256:36bd6acc3b2765a06993d06d8ec5ecd1ca9b5605ad21ae182dfb93459cdb3c47

Observation 10aa9ec4-ba0e-43c5-90e6-19807c442f98 · outbound

This paper cites s1: Simple test-time scaling,.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs s1: Simple test-time scaling,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.842800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.842800Z digest=sha256:3c9bc14f757dd8edfb1645e454a29fc55f5c7b6605a8c1facb4aa48f2f4b80a2

Observation 52911944-a89a-47e7-ae74-3accc39b8d79 · outbound

This paper cites EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.952530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.952530Z digest=sha256:ca3841bdfbdd7a56bac8501240f420f60d7e6c44f443bd9c7936cf1a6161e9d3

Observation c22a6a3f-f6ff-4d52-b80b-d5f983955527 · outbound

This paper cites s1: Simple test-time scaling.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs s1: Simple test-time scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.907112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.907112Z digest=sha256:66e8ee046932f3a0dee543e7cef5fff58d417b607a5db97d721b9a84fa949a60

Observation a5ee57e0-46a5-4534-b3f4-40b377ee033f · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Maximizing Confidence Alone Improves Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.075191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.075191Z digest=sha256:bec04195f900a5f5881a094134feef06692e2885171038c730299e295da9de8b

Observation 7788e3dc-c34d-47d8-9433-c4febbdd1c91 · outbound

This paper cites OpenAI o1 System Card.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs OpenAI o1 System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.026991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.026991Z digest=sha256:7677834941799e80d7e3017e398434c8d5969d6dfd755463d94a77eb3e869a5c

Observation 9501d6f4-81d4-425f-9b26-ef0835d084cd · outbound

This paper cites Optimizing anytime reasoning via budget relative policy optimization.arXiv preprint arXiv:2505.13438, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing anytime reasoning via budget relative policy optimization.arXiv preprint arXiv:2505.13438, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.114011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.114011Z digest=sha256:0d7cdd57b876e8538150e0140356f51155eb7aac506fb7093d67cc32509c0155

Observation 82baf72a-03f8-433d-926d-851f97505323 · outbound

This paper cites John Wiley & Sons, Inc., 1994.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs John Wiley & Sons, Inc., 1994

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.081368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.081368Z digest=sha256:38b862636644e968cf5f3f58587ce0bbfd2421d8ee70afcd9bff5754cc82ce8b

Observation 83245270-4548-4a11-8767-ddf00921963f · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.410502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.410502Z digest=sha256:a74efbcfddd60bb1aeaa1ec825d1d36c1371511b0733e56a9201f7deeb0f28d3

Observation 93881a94-72c1-4b83-815f-b6aa2f1a469a · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.330777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.330777Z digest=sha256:4401e0c6b539271634422295dcf6ebaf4cbc1499ec21724376267fa591aa0712

Observation 1782b1e6-2e5d-4e88-bd43-c2cc14ef8572 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.538173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.538173Z digest=sha256:621e4fffacbda7ee7eb283894590e2be8be81c57e43557b2b25567d7f83496ad

Observation 7082dd8b-ab49-4565-b43e-4b6fa9fd41ce · outbound

This paper cites Proximal Policy Optimization Algorithms.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.481185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.481185Z digest=sha256:418fa2ebd0e7728f8205563d68211084ab2347b85f89fe9d46822f85065d6da6

Observation d155d0d9-b405-46c5-9231-8f57d771e0eb · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.653994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.653994Z digest=sha256:c698aa785be677c9aa57b4311d600c22ac4fdc9b2bde924f57946ed8a4796dac

Observation 8ba954c4-8f3f-4ec8-8b34-bcad6404e7b1 · outbound

This paper cites Opti- mizing llm test-time compute involves solving a meta-rl problem.https://blog.ml.cmu.edu/,.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Opti- mizing llm test-time compute involves solving a meta-rl problem.https://blog.ml.cmu.edu/,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.601303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.601303Z digest=sha256:9ee3ce3d60b2aaa720089eda9fdb818ab0939fd935c808472e6a57d9debd3dd3

Observation 93354397-44c3-457a-8c2b-ad6a5463736c · outbound

This paper cites Spurious rewards: Rethinking training signals in rlvr, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Spurious rewards: Rethinking training signals in rlvr, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.887825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.887825Z digest=sha256:30a15aea402609307d1c14b6acf734b1c8f2e7a0c1b88f9307c5ff8b96b0e4f8

Observation aed5a159-d96f-420e-8049-51292198c6ff · outbound

This paper cites Can large reasoning models self-train?, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Can large reasoning models self-train?, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.763477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.763477Z digest=sha256:91c7403deaa2e144c4d7919d80215523d85f3771310ac709b45404e749f7d0b9

Observation f28761a3-3781-4ea6-ba56-de81838ac795 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.093876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.093876Z digest=sha256:8f7f0d3e854006c5c33252c376012c6c58ee744fbd279e1d2bf7b4cc15dec9a2

Observation d72ebdd1-a06d-45cf-ae99-5f5e48a5f708 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.977837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.977837Z digest=sha256:539bb355945f1f48124b3574909d6163aa8a39d622b9c30b501ec61453848226

Observation 3773da43-37af-48b7-9350-02da630331d1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.325951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.325951Z digest=sha256:8fc41b0a95daad2003f9328be382f809154bd644d7bfaf7209def28729b66976

Observation 207ae8f2-58cf-4814-a505-d16add4c7125 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.240528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.240528Z digest=sha256:43d11aef2d61559a2c7a2af227cec7f562671fa3a3862669021f050fbd70a5b4

Observation d722aab7-75a3-4e2c-b6c7-a7e5fb01397f · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.528670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.528670Z digest=sha256:35ca963ab68cf5754b19d50ba6e45b932757c93cb8ae4a1622bd9164599ac26d

Observation b760d7ee-a2d3-4a3e-a999-b4747dad6391 · outbound

This paper cites Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.400160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.400160Z digest=sha256:f5a430261efa278042f6985868a537d08143f4badfc39c3510fc55701ba606f9

Observation 41efc73e-f3a3-4f0f-a89c-97b4007aa361 · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.718332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.718332Z digest=sha256:7acc41491dee73fc78c145bbcbf51607858977aabddebf4e40b4125bd80d915a

Observation 014f4051-8c67-462c-929d-4f6ef5cb1548 · outbound

This paper cites All roads lead to likelihood: The value of reinforcement learning in fine-tuning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs All roads lead to likelihood: The value of reinforcement learning in fine-tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.616812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.616812Z digest=sha256:4504e21d70f2a53450f2ff431e4478cd32cd69540850c2d3b2eb2ff07adca757

Observation 741f74ff-94a1-47c6-a190-cdf584061006 · outbound

This paper cites Open Thoughts.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Open Thoughts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.890910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.890910Z digest=sha256:c1ae9bdec2d1e0619e9f99265ba3dd9c9fc33a964735cd647972919781f6c4d8

Observation 878ef22e-f7fb-4336-9912-49ecb9b789eb · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.800270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.800270Z digest=sha256:be5fbd56a45d545fd5027819977cb7380aefed7aa48eebb78516c3c6602c5b0a

Observation f835cd85-0c80-4003-b734-4115945c1f75 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.161726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.161726Z digest=sha256:d4d9e13aa2edeaa27c6272456fab116131ab9ead05b137f87a768ce513a26c4c

Observation 86ac5cef-66ee-4c4d-9405-8c67ec8399ff · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.029441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.029441Z digest=sha256:c6567e367b8677e8d9d7875a7da9947505747c7a9caa50adebf975abe8441ddb

Observation f0eaa6fe-7121-44d7-a84b-6ce3418d5c79 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.367441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.367441Z digest=sha256:6c94d65f1abd72550c1d5ac8698c3e25b969d371d1d85d7a247af809a1de5d98

Observation da1e103d-0c2b-46f9-9ec7-68d171daa5b5 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.284245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.284245Z digest=sha256:9b278e3f2b15781ace2cca67791db7d191dd90369065d997b934809fd0fc2ba4

Observation 9f9f2196-f3c0-4924-9bb5-06e05a112b7d · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.615178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.615178Z digest=sha256:6ffc3e0102c998e01d0ca8ce5a108f962e14cca0e8f5c3dbd708273195bc86f1

Observation 968c06be-1c7f-4e33-a8d9-a2e465a1f2d2 · outbound

This paper cites Qwen3 Technical Report.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Qwen3 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.500826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.500826Z digest=sha256:98d95850588478c3f91dd18aa0c727b57756d1ebe56a74135b4f7b84c3115d22

Observation 8b61a26d-890a-4211-83ab-ee5887694417 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.861106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.861106Z digest=sha256:533147581f88551b897864572551b85cdcb01e69a29b05d6888c689b7f64e121

Observation fbe2ba65-17b0-4fc9-a82c-421acd9a1363 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.742448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.742448Z digest=sha256:31cefbe1f9ee0a3f9adb6a911e30f23a8cebe33d9a6454eb78a2cbdd100a92ce

Observation 26d30ab5-fa15-4fc6-bc30-355c07bb69e5 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.047432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.047432Z digest=sha256:3d5ffd67603fb8be942295b77bd42d1850710c6dfe925fa265190212a98f175b

Observation 233747cc-2482-4e80-a50b-534029d8f11f · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.931120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.931120Z digest=sha256:05cc04693e079f3f6a31123fbc92b5fa66ab052518025b712ed0a51fc0c62ea2

Observation 5f7f6153-5475-49de-8a4c-b66f7025612a · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Star: Bootstrapping reasoning with reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.153724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.153724Z digest=sha256:43360d64d904f5f6a77d73bd0abea467cfa16f196024c143ee34ab7a7921f51a

Observation 5d7413e4-ec3c-41e5-866f-878829bce260 · outbound

This paper cites Learning to Reason without External Rewards.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Learning to Reason without External Rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.482279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.482279Z digest=sha256:2170af60ba6a372dce4b9a87db6506f9edc9ff835c667e1e4d9f567dc22369b1

Observation 6d5d42ea-cccc-42f2-9151-fa579a2a658d · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.353983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.353983Z digest=sha256:96cabef53af65200757447b710dd893c4294010f0826440ade93b2d72a0804ab

Observation d8b24c0e-a7a5-4ddb-bdfa-0e4246605f7a · outbound

This paper cites Looking at the other numbers: 77 - 70 = 7 97 - 73 = 24 (interesting, we already have.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Looking at the other numbers: 77 - 70 = 7 97 - 73 = 24 (interesting, we already have

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:34.993578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:33.298133Z digest=sha256:76b7bb428fa25bce21866221387587fb450fc22e9562593afb641f2024521360

Observation 73793d47-d65d-4c6b-be67-7ad5aa220217 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:40.735256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:30.603388Z digest=sha256:fc6859d7944f238b4c3e91a53e34343d12ff114f08c283503ebaa6830ef90623

Observation 16ec9fdc-ddb2-4d41-8eec-2bc04cbf672b · outbound

This paper cites We need to get from 37 to 466, which means we need to multiply by 12.5.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs We need to get from 37 to 466, which means we need to multiply by 12.5

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:40.511892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:30.705279Z digest=sha256:5b84839a90a86ca99470a7439e2570a7b5fd0cf8db24caac08044ccbc6e6430c

Observation cba0c388-9ba8-48ee-91d6-0e00a7c42aaa · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:40.298433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:30.826305Z digest=sha256:88a572171878a3a41e99f87407c9c74daa170dc831ebfc1070b062c880f0d353

Observation 7703ae3e-ae22-45dc-a9ec-624e83181129 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:40.090080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:30.947729Z digest=sha256:bed72bf918a7f4ffc334ba04102839852e6c7908e1adc5d5e5ff43289f2906e7

Observation 38c15b0c-bec0-4193-848a-65acd990faf9 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.898527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.064851Z digest=sha256:e59c8e7012c258b0947a294abfbf2e2b9dd90f228182fe8af3342d617e840fb3

Observation 2b8878db-da0e-463d-b3f2-5872b395b83c · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.646657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.157389Z digest=sha256:dd6ba9f485e7728bad353a2f3e5967933225d8a94209ee91ecb1080eb56d3878

Observation 48455646-a978-45c3-8a4c-22f364a55508 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.438320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.247123Z digest=sha256:39a433e856884a4f29235f4df7f8cf4bbb779667ff0f654ed5e62b3032d0b528

Observation bcb27548-26d4-4336-8b9f-2c7770a86ee7 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.216261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.382722Z digest=sha256:f51f6cbac69317420b8750b227e31a209f06c6098384adf2d261ccf05f7caf3a

Observation a42f1119-c124-4ced-840f-64137e979db0 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.009793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.517583Z digest=sha256:4db1b18f0c97b9ed6a81c1ca8257f07ea0d0545a77ef098b86385c5b70c4fa9f

Observation 66d47b82-d388-4183-a294-7b9d7472290d · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.877821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.600499Z digest=sha256:0cae1e8d516fef36ad832ebb041b6783ae0d3425d15bc773831fe68158525c9b

Observation 476391ed-9599-46d5-ac75-e69fa5892a14 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.657484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.690026Z digest=sha256:9ca95799756576a504abf8fb314586753f9f65911003e9735853299f17de193b

Observation 7f1d463a-e987-489c-99e0-3e57a75edd6e · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.408750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.769004Z digest=sha256:475201ae00e3e698df8fab4240914e3dfbc0336e245824cfeaa067580594de8f

Observation 75ba2295-f46d-4876-a35b-e71b4b53bfff · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.133653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:31.928589Z digest=sha256:ec68ba2d1f16c35b91e26a149df2468aa75cc68d434b86b58df4c365a1697c1d

Observation 66dfd81e-85ea-4ea7-bdde-1fa303f6a108 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.931265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.068744Z digest=sha256:e0a9e423df7ff5acc74b08a197a24d5981f9b53ecd85fe0740d20203ecea9331

Observation b2e2a2a8-3ad0-47ab-9d10-8fe797d320e0 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.716974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.154019Z digest=sha256:63ac30cf27fb9af66ed2d4ac05d17d2eb05ceca347e323bd0837ee707b0de4ad

Observation 884f3572-8433-4f42-9e4b-22f75dd0c333 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.520483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.248190Z digest=sha256:31c1d13e23df1eca347183959489584f5755ee66978ffc833d72769eaf5f91a1

Observation 437b83b2-677a-4490-8b68-89b6d5c1cd3c · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.291156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.343942Z digest=sha256:cae8576492b39d4b191daaf7711766a3321f5eb198609ff0a0d5dfb192612877

Observation 269eedb9-3091-40f7-8272-7aefe649550e · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.088649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.446315Z digest=sha256:745668c92bacf64b4d615309597773d3d8def407adad4e2e9915aa11a0448e11

Observation 3c5a907d-1706-4adc-88f4-31f4dc9dbfbd · outbound

This paper cites 31.5 (not helpful).

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs 31.5 (not helpful)

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:36.816490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.586951Z digest=sha256:955d0e6d72921bbdf0b2f425f251a16cfcd499196f6399b106230bde492680dc

Observation 7db092d6-5cdf-4905-bf33-b562e09ffb1b · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:36.569151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.661655Z digest=sha256:c0eb1817da8647b0f1f535fb84f469765b2597f0c077f78d73f29e3c4b80a0bc

Observation cd6f2c5d-3087-473a-85c4-ca16269bc216 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:36.358782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.768424Z digest=sha256:f1d4ca7a5540f2f9b11a78a4a1e9479025f7c6a60364aff904844ee0bfc81d47

Observation 8fa66cf5-8f1c-40b3-b0a2-9e5cf583b650 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:36.118343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:32.882107Z digest=sha256:d8448d4d6e3908554a778cd7bbbb9caa18f6b77e62fdf3185d5bacd3936c8778

Observation 2efbcaa7-2bed-4693-941a-5eb70b13dbd1 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:35.933890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:33.007560Z digest=sha256:8b9e4a610856c649aeb8dbc2f465ee372f865d98cd9460c01658c46e30b5c61f

Observation 415a2322-a148-4db3-849e-c71a98eef83b · outbound

This paper cites Hmm, let me think about how to approach this.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Hmm, let me think about how to approach this

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:35.609076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:33.139220Z digest=sha256:4d59a8e2e88296dbf09a17ce10d380083b49f5be603a8cf747bacbbfbca8870a

Observation 9759744a-aa27-460a-9737-6bad422b14a8 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:34.743229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:33.393958Z digest=sha256:931e92040f58fec99022f188e71ecbe944b3d9f3b3e2f28603e39f869e5d78f4

Observation e8a353ca-341f-4290-925d-faf4647054df · outbound

This paper cites Again, multiply 347 by 5 and add two zeros.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Again, multiply 347 by 5 and add two zeros

Reference 500

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:35.287512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:03:33.212215Z digest=sha256:ac289f77ab564c83c7fb62140c49260c6eb7416f6375e47d48f6ab55e38deb98

Observation 7ca58a45-cbde-464b-9c44-32b754e36ef5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.469591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.469591Z digest=sha256:70bff0cdc6a5a079ea845927f181294c7e37e670b9a81d9c721f1bfe57627814

Pith citing papers

Observation 9749439b-2f45-4577-a302-9162374af12e · inbound

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context cites this paper.

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:41.937333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:21:41.937333Z digest=sha256:9f0e0bba57b1d6493277a6996382aef1ff8ff80cdef7b50e0adb76eebd488f98

Observation 2dc96f34-873c-4b75-ae35-482082edbde9 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.163389Z digest=sha256:c1d55cae7811db47f4e125c7e47810a71890a9e224f2f553e306217217471d04

Observation 62e79bd9-5de4-4b91-b579-44a002b56daf · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.534182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.534182Z digest=sha256:19ff1ea1da1a0f9a1fe4f4ee61f1122f61b7e2e36b102ae26a145d0d626bf6f6

Observation 6a1e0385-5950-46db-a390-5826abdc1213 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.509893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.509893Z digest=sha256:7d804d5c2ea5f6de81f2301e6852985f8c26b78e124427859f42173b002b4cb6

Observation ac25798b-d911-4be4-9de9-432a70e14263 · inbound

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning cites this paper.

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:27:44.448791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T10:26:45.961575Z digest=sha256:5eb3f16c0f1107526d95ea4ec86ccad53b3c06448dfb161726af4903a809dc12

Observation 96a65a17-114f-4a3b-9585-6ae340cd7f01 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.923996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:74c19fe445a4e803209d79de1fb910f3d43cf0b8ee6892cd3f6a459d1eef6cab

Observation 97b7fe68-7dfb-4165-bc4e-0cade0795304 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.402466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:35197a034f7975171c9e22852a4d7a32a640da34afc209545e5c971f39922d22

Observation 11b5d117-b3eb-4c77-8be0-13e8c12225c7 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.517079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:0826ec0b2efb5eb6f6494183143e95c39a02718c414d1bffd0eae3b060d068b3

Observation 9de63b67-d0db-426d-ae15-87987fce5a98 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.652333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:972fa6669cedc468576d5470e105bc0cf13d21a9cf763c657b3ca9fb0ceca7a1

Observation 978bf51c-3054-476a-bfff-6e3746dd43dd · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.433620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:0310bb01b7745cdaea1920a6f9b9d7682a16e43ea12616e9fd2a0fb6b7c55acc

Observation 18fc5695-e22c-4d53-9796-fff3c94285ae · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.964580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:cbb02c9e0d0eda2259ddd8506822d43b99a95b9d949f5eea7c9f64d2b5f6f6e6

Observation a38eab6a-6954-4594-88ba-7018c0acc17e · inbound

Parameter Exploration for RLVR via Variational Learning cites this paper.

Parameter Exploration for RLVR via Variational Learning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:13.013301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:13.013301Z digest=sha256:e57eb75a20a6317642a36f97d12c4ecaeb33fb22f60fbec8906e6f67658d5873