Pith. sign in

Paper Citation Record · LEDGER

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

As of 12 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 45 inbound Pith citation observations for arXiv:2508.10751.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10751 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:21:08.269218Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:16.232644Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:08.457909Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20fe78a0-c3b7-4dba-9e54-00d8c30a84d5 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.023052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.023052Z digest=sha256:13424e586a21d1807781721ed9ed033dabb558288358b3c34ff6b78f3375682c

Observation e4f65368-aa33-417e-a8cc-39db8d830f48 · outbound

This paper cites Aime2024, 2024.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Aime2024, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.639070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:04.149401Z digest=sha256:725feed6de61e0502295e011261554c7be48ff3a80bd0724f232288fb7dd2fe1

Observation 311c3a56-09ae-41db-bcc8-c81911226bee · outbound

This paper cites Aime2025, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Aime2025, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.624512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:04.271095Z digest=sha256:af31a5b7cda834472f5bd130772f53360719acea13e2d393f2ac1fd380a1acc6

Observation 62f72c5a-6235-4654-bc26-fe632c7dce51 · outbound

This paper cites A Survey of Exploration Methods in Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey of Exploration Methods in Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.419668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.419668Z digest=sha256:103f60bd537eb1880fae401c802d3bc043260060ff84405f4588a911d1cd29b6

Observation e7c1322f-379a-4b7f-a5f3-bdd4942d70a7 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.577447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.577447Z digest=sha256:bff0d682f94556c2a639990763b541f90a9117a467b59474aa9929887c0e6945

Observation 3343a24b-c34c-4c73-b8dd-5dcf52e709bf · outbound

This paper cites an unresolved cited work.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:21:09.609052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:04.749033Z digest=sha256:823e23642ada6310c2b8d8d6d308cc23be084ec3e67d8511f8335c6fcd11db51

Observation fd4291e8-8156-4eed-a22b-59ee16b57e10 · outbound

This paper cites Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.915607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.915607Z digest=sha256:b7eee74acbaa8ec5403782f2d2896c3eae17c9ab6cfa60f2471619719ec69d51

Observation 42cc8885-3790-443d-b3e1-dab8a68371a9 · outbound

This paper cites Improving large language models via fine-grained reinforcement learning with minimum editing constraint.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Improving large language models via fine-grained reinforcement learning with minimum editing constraint

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.593888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:05.110565Z digest=sha256:55308d0ebe7996247b13c639fd1327f71996251680f60f019726acca739f25c2

Observation 9a097a0b-09ff-4ebf-8733-eafab4a5893d · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.350521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.350521Z digest=sha256:b97cff4d66505793e7ebc15a0f52a3dc4ee46e90a3c4fa190012af2bbe867f08

Observation 99805c56-6e20-4578-aacb-2b4bb2e1fcf4 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Reasoning with Exploration: An Entropy Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.510301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.510301Z digest=sha256:b0714f0a807c091f5117a664c8080be2245aa7ed126fe90b19cf671df518745b

Observation b7a6e9fb-9cdc-4e8f-81d3-5e95ae19aa19 · outbound

This paper cites Thinker: Learning to think fast and slow.CoRR, abs/2505.21097, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Thinker: Learning to think fast and slow.CoRR, abs/2505.21097, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.724526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.724526Z digest=sha256:106486c76a01e49e90a59e25dc7bac8f9aee91202647fe4ba6c367a9b6c53a86

Observation 1139fa17-a670-4e6a-8954-5c8efaa7a883 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.883163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.883163Z digest=sha256:f0f9faf59730447bf3b8f711d983334ef952cbb79edad1c5a13943552a08c079

Observation 2d842488-652c-4c9c-8f1b-b8d7477819f7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.013467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.013467Z digest=sha256:2d0a37df566bd5f344cc327c5ee768295f113e0f57dbe8a9581d6c0505ae387e

Observation 6c23da36-7ce8-4391-813e-3aa311c32bc2 · outbound

This paper cites The Llama 3 Herd of Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.145355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.145355Z digest=sha256:85236cfe40486f211bd73d58fc27997a5c3796d9c2a26974e0b3dfc32f8ed5a8

Observation c0c2e054-71b3-4084-bd1a-8c9069851b83 · outbound

This paper cites Stochastic first- and zeroth-order methods for nonconvex stochastic program- ming.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Stochastic first- and zeroth-order methods for nonconvex stochastic program- ming

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.579100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:06.230165Z digest=sha256:008359419b498862cf6ac7e58dedac4048d60f66cb22636a79b930e375fba884

Observation 7084e6e2-2365-41dd-a044-67fafd312f2b · outbound

This paper cites Bootstrap resampling methods: something for nothing?The Annals of thoracic surgery, 77(4):1142–1144, 2004.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Bootstrap resampling methods: something for nothing?The Annals of thoracic surgery, 77(4):1142–1144, 2004

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.563898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:06.330232Z digest=sha256:6964dd25af72c2b0758dc1e322ebf6060342e62fa0ca95096a7f5b687e04868f

Observation baf71de0-f6fb-4fa1-b7ac-c21f72141e0a · outbound

This paper cites Seed1.5-VL Technical Report.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Seed1.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.466198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.466198Z digest=sha256:ba922ae02ef28789a7cb7f8d5f7eb4a256e0aa809abf238e5b34a00b1c6901d3

Observation 628887d4-4e15-42f6-bfb4-b9b8d72704ec · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Skywork Open Reasoner 1 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.606460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.606460Z digest=sha256:3e48d316be73358b3a395c490336045dd38577311e33dea518cb4715b92c077b

Observation 245b89ef-4157-44b5-b6b4-b08d30903107 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.704021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.704021Z digest=sha256:ce4340c3350943324430592d43a49a4b28984c26226a9287a14d6583bfef629f

Observation ac01cefd-0396-434f-892a-51412916c5c1 · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.806080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.806080Z digest=sha256:6bd9a12969a149d8dfcc4ea24b16408e916618bd8aea19898af72d541aa879fa

Observation 2e500108-796a-4d7f-aa74-1553e62ba1c5 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.909124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.909124Z digest=sha256:2047cd0ee6c47e9d46c688b5438ef12f1be2eb6b372b583132d483a848512aea

Observation 6ee29c13-b5f9-431d-a2c7-d3841c64ad5f · outbound

This paper cites Test-Time Learning for Large Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Test-Time Learning for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.016294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.016294Z digest=sha256:eb25e382de3509d2969799319b46ec1b65f57045955c375346d6917b26a799e8

Observation 3c53f383-d82f-45cc-81e8-90294d83ced7 · outbound

This paper cites OpenAI o1 System Card.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.122698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.122698Z digest=sha256:e0f806681ba248a59cad9642b7c951644459f84b1ccc87948ae1aea939764ec0

Observation b84b9ea5-a56b-42ca-a579-831457015b47 · outbound

This paper cites Mistral 7B.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.228054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.228054Z digest=sha256:02ad5bd2c769048795c6005b42d523b32653bc574afe95c2611492537103f340

Observation 280a78ad-13ba-4915-9c95-3bed5a25b21d · outbound

This paper cites PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.329389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.329389Z digest=sha256:2610d75bb80ad8901bacf666a9cc176f604364108f627e3148dac783e1041df9

Observation 51ff1928-5d33-47cd-8da2-fce57a9ac1d9 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.431738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.431738Z digest=sha256:9d2116671dd9e095ddc5f17e6befbbad554bb9f2ca2c0d2f045ff7273ea22248

Observation 7e909de1-669d-4dfc-90ad-4dee4fa37e3b · outbound

This paper cites PP-PG: combining parameter perturbation with policy gradient methods for effective and efficient explorations in deep reinforcement learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models PP-PG: combining parameter perturbation with policy gradient methods for effective and efficient explorations in deep reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.548074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:07.527540Z digest=sha256:bcec69cc83d82ef53cc0b633bdcec60bb9b86637cf8ea09183d28345b8de9171

Observation de686465-cf53-4b9e-89c3-2a3fdf6af3de · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.634233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.634233Z digest=sha256:a46309869c644c126eca1b878211382d4b441b7b53276af41dd5fe29d111ebdd

Observation 31165b7e-a287-40f2-af57-a8ecee7d2dbb · outbound

This paper cites Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.730797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.730797Z digest=sha256:d189feb5d0d580c6f4b0f4b81584899afbb9c804c730746993625d8571c19475

Observation d0e5f048-ce60-4b9d-8933-0fe3d1f6116c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.843467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.843467Z digest=sha256:d7ae24797a76fc2eb263da841dd01f498ac04c965dde097c0040efd7d35c5899

Observation 135cdd04-f9ce-4df2-a4c2-2a1b862eefb8 · outbound

This paper cites Inference-time scaling for generalist reward modeling.CoRR, abs/2504.02495, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Inference-time scaling for generalist reward modeling.CoRR, abs/2504.02495, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.944339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.944339Z digest=sha256:c38eacfba4e233dd2ac504a55cda3b9366b7385dee47daf470b8b58597cbe608

Observation 63245719-83c1-421f-965f-4e7e4baeb7c0 · outbound

This paper cites Learning from Peers in Reasoning Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Learning from Peers in Reasoning Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.060154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.060154Z digest=sha256:4cccdbbce96fb77964668f7397e6f0a86f8eae554a9b1e0371b80a2cb91c77bb

Observation 9eaa3cc7-49dd-4760-9e89-0b0c97512861 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.144794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.144794Z digest=sha256:327488c83980d1ed96ec7cb52c63836d71a48d2310c92cedf535d18bc7aed587

Observation 21f9c201-1872-4147-bbf2-d17c4396691d · outbound

This paper cites Nesterov and Vladimir G.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Nesterov and Vladimir G

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.532062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.150455Z digest=sha256:b9bdeed44b41b945c28a6c567eb847a3df620bfeb62ffffad80a8b963b675e43

Observation 4c3ddb77-da42-4ce6-a580-5cc87dd8b440 · outbound

This paper cites Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T20:21:08.865359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.155057Z digest=sha256:ebf9841094f531f6b5d13162b8453315db1251948badadf82576f7d9515aaa0f

Observation 66f0cbd1-9156-4f7c-be0c-512a26e701f1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.159316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.159316Z digest=sha256:4c002da15af6e623e74700caa6c224f01ad295a69b4a01f2fb0586f97149c93b

Observation 2dc96f34-873c-4b75-ae35-482082edbde9 · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.163389Z digest=sha256:8e3cdffdf83c676320a5f196f73ed90f318e8ff4b3864ee036ef12680620f751

Observation 06d7b561-f235-48b0-80e6-e4599b64f9c2 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.167577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.167577Z digest=sha256:b6a57d8554aab46c33d956d6e8bbdf040bcafbc5fca1d3ec084a3213eba992d7

Observation 0d6b4713-7f46-4386-95c7-6c6a7dfa789f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.171925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.171925Z digest=sha256:7e9bc85b641f85a8b3bb25f49a1c932afde1a0eaa809179a796de9dddd7d0581

Observation 02db5137-7dcd-4074-b5f8-5b25a7ef2719 · outbound

This paper cites Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.514447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.176068Z digest=sha256:b74d352d033ff843517b8f31b80a1a98fe23d14a3e0044dcb989866b7c58af7e

Observation 871c4546-f675-4c08-b313-9a27eaa3a600 · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.180531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.180531Z digest=sha256:311226bd71d320dd5cbf2663a3d10e90f0043c9adeeff6722e26eac59d73a13f

Observation 8e997e85-8055-440b-a0ec-1515799f464c · outbound

This paper cites Optimizing Language Models for Inference Time Objectives using Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.184900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.184900Z digest=sha256:afc5256180facc36fabc4f136c9e86528ed9fd4e473b9753c5a8d7894e72b6c3

Observation 628a0a5e-3563-4f67-926e-7662a64adb4c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.189939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.189939Z digest=sha256:abceada124b6dfa9d4a69761c908de205adad5b6517a9625dcd6d1b399243f3e

Observation 0d71cc55-8d76-47da-a7e8-3bd0ded1ff70 · outbound

This paper cites Reft: Reasoning with reinforced fine-tuning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Reft: Reasoning with reinforced fine-tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.495740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.194262Z digest=sha256:7034915a8f31662d157a2f79c56d26c102d87a1f265d2e1005e2c65d70fa6971

Observation 96da9560-e314-4886-91b2-7cde506d392d · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.198160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.198160Z digest=sha256:ffed536a25073a429a281347c42e6dc0f7073c51c647bb9a7b87c465d338cb5d

Observation 46f7412a-5d75-498b-b50c-d766cf13e929 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.479040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.202417Z digest=sha256:deb3c9cf16c2c0867ed111457b25d425a45fb75be51a62c78f68a2fcecf6d0b9

Observation 8c9b5026-2071-446b-9b77-2a7369af0836 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.206755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.206755Z digest=sha256:f96a8cd993207190d4d7a2304f7968d8922f69e9c169a1b90536ce54c45d74a4

Observation 8bbbfd40-7616-42d9-8616-39e8c6566b31 · outbound

This paper cites Williams.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Williams

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.462263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.211104Z digest=sha256:99733734558352dfb77cf779eb0e80f4a4d0bb32f6582ef1670fa6e15da6bb22

Observation 3e5fc915-e57f-440b-96a8-326d3280e5a2 · outbound

This paper cites ARM: adaptive reasoning model.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ARM: adaptive reasoning model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.215587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.215587Z digest=sha256:a692d1de45e2fcdce9600008b205d59d00a047cccf355c8adf50f34473b9ea72

Observation 2b52b9ed-e82c-42e1-bdb1-a84c211552fe · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.221916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.221916Z digest=sha256:ca5eb2b3f17301f924d31a398c44a5622a56fcf509fce2817c89150204c2ed13

Observation e394acda-e293-45be-98a7-f889ebaef45f · outbound

This paper cites Qwen2.5 Technical Report.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.226479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.226479Z digest=sha256:9abeca65b478ccc2cba7953faf47e01f02d5382021f3ea07c1b03e83638b140a

Observation fe7ddb09-536d-440d-9862-378a9c939715 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.230632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.230632Z digest=sha256:d462c1a486fd6ccfb1a892ff0f30c1a01cb26f80c3400ffebb4fe6eb41ef823a

Observation 69c396b7-6c36-4336-b590-bd8dffcba871 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.234662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.234662Z digest=sha256:9ee6267945a919df4530763443c11547fbd43ce789899790b2bd3df8490c1398

Observation de3cb2e2-c7c8-4fea-b012-22ad3ae2bae9 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.239014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.239014Z digest=sha256:c0ba510e1d8e2ce911fdc2530edc762c8168db6839c5162269e905e910b27130

Observation 8d6ef5ab-efdf-4d0c-b7e4-1c4240bdf064 · outbound

This paper cites Boning, and Dina Katabi.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Boning, and Dina Katabi

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.243328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.243328Z digest=sha256:4696b90244905a84ab14727bbca2b6b2cd91ca88d7265863e618b8638a46c3f2

Observation dcac152f-aad7-4d74-a694-d4bd78ebca09 · outbound

This paper cites Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.445445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.247207Z digest=sha256:07ec3cb1c0cd888204325fb7ba43a28596730c4f9cdd1b7840da364f4f9bea9b

Observation 635baa46-0aa2-4c0c-83e7-d034e1aa98a7 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.251549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.251549Z digest=sha256:7dfbdbedbca70c3d44338b325a1ddf628d665f4f737a1647e59fdfc8240705fd

Observation af82a7c5-9336-4f8c-a046-38fcc0f9ade4 · outbound

This paper cites OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.255765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.255765Z digest=sha256:821f1231256b22d6ae93accb9f28ada2373ba78ac01c8d0aefb47746c3a32171

Observation b7303c30-23fe-4d3d-afb2-afc9e91ecb71 · outbound

This paper cites The surprising effectiveness of negative reinforcement in LLM reasoning.CoRR, abs/2506.01347, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The surprising effectiveness of negative reinforcement in LLM reasoning.CoRR, abs/2506.01347, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.260188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.260188Z digest=sha256:c4f15f8a446fae346479a4a2e13f9e14548af0a81f79174ecad335ca06daf9f2

Observation daf99f1d-3ef1-47e9-b666-abb418f71612 · outbound

This paper cites an unresolved cited work.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:21:09.429269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T20:21:08.264276Z digest=sha256:e705595afc762c1dc910d2687fb57b673e8515eeaad5434e76643c4ce69a7e0e

Observation fac393c8-196a-40fc-aff7-6f40df93604c · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models TTRL: Test-Time Reinforcement Learning

Reference 61

Resolution
malformed identifier
no resolver link, observed 2026-08-05T20:21:08.269218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.269218Z digest=sha256:4b5c1687ac25cf26ad4c7500f0f3b80e375386415c0d75c820d6e2e09b7b991e

Pith citing papers

Observation c40f1373-96d7-47d6-bf88-a85fa7fb7c7d · inbound

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR cites this paper.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.124820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.124820Z digest=sha256:f670ea598b20ad03392eb64af5f297fc3faee5c549b34318b7a139022c838b37

Observation 799af097-4862-4a75-989d-849b0b503152 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.591192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6dc4f8761b888fed0d2eded2e167e07ac1284c336423b53cd40f8e9eb2bf432e

Observation d36f566e-f01d-46dc-9179-90ccd70433a4 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.473018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.473018Z digest=sha256:08eda5820448eb0d0471f35368cd3a0f3bfc19a91261e4f6bc2f714aabb63ea3

Observation 45b8b181-2879-4fdd-9eca-4e0f414cdbfa · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.160513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:3b135857cfa3bf8ab44f6e0036835737b2ee0f31170587c1fca80face4afb210

Observation 88eb6109-4516-431f-a947-b6785bebb238 · inbound

Emergent Slow Thinking in LLMs as Inverse Tree Freezing cites this paper.

Emergent Slow Thinking in LLMs as Inverse Tree Freezing Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:46:24.288452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T12:43:49.628082Z digest=sha256:d6db13bd68fefcd0ea5848c59afbbf171d63709055ed61a3c916ff123f0cabb0

Observation 191f6cdb-d00d-4350-822a-e95904af13b6 · inbound

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs cites this paper.

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:33:56.981364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:33:56.981364Z digest=sha256:984d45c138ab9958fe943cc1523ff4cc1aacca5cd0191257e7ccb17b50570b6c

Observation ce897c27-c9b9-450c-adff-e09bdc6df006 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:06.058912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:06.058912Z digest=sha256:347d7501d45f63b8c1052ac4598716fdc9bf4e854803338ed5220e409b4ababe

Observation f0672f5f-5a20-4c0a-a0c0-18135e3042d1 · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:22:31.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:a437554a75adadb58a84af08cc6da44fbd1e21b9a93dc244a08210fc0151947e

Observation 365887bf-2e01-45cf-b80d-f1c14d8774aa · inbound

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning cites this paper.

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:44:43.950446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:44:43.950446Z digest=sha256:a363cd0f2335478fe446bd1d004a67a9ebfd737455b7be2a5d66f031532df257

Observation ad4def12-7ced-403f-bfb8-dd4bfafe21ec · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:41:03.928079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:e275d0d7f6cf45d66227019975400c074b6288928aad961b4dde7ab7562500d7

Observation 493c6899-d3c8-491a-8f31-7eff9bddad66 · inbound

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models cites this paper.

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.810362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T07:05:45.081955Z digest=sha256:98cc9dabb3d1882f081e179c8730652e600e298349ed1d93717ad8d01c175119

Observation ffb3f9a7-bae6-4387-a39d-f03a74925aa4 · inbound

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models cites this paper.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.470700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:7cdd5ec50a59232c788489f11afcdf45ec5076103c278dc3cd4d45eb0a95457e

Observation 6f9a438e-d767-480b-9e90-66e45d68f524 · inbound

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity cites this paper.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:15.655822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:5a8b5f517c8524a40e04a7d66f0cafa3d527885b23bf0b76774cae7abe36da2e

Observation 3a3e0fe9-f636-46dd-a3e8-249b9504625c · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:46.295094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:e221364daf6c50673deef519a6ad55a24a53cbd320aec70b98f813bd73fd4c7b

Observation 51162d63-f47d-4295-aad7-d13027ce7f2f · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:56.520021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:94fa156ad7149dceda3857989a769dcbb33e0be91be0f4a87569c27c759fb50a

Observation 1b4bac14-8d4b-4e1b-99f9-62665a56480d · inbound

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR cites this paper.

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.050125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T12:40:53.063991Z digest=sha256:b26d1cf69008781c4d0ee37f16b256c47359b55fbf0407165edfc2b9a7d5292e

Observation daa2a35e-d39a-4ae5-b48f-d4d49bffec98 · inbound

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents cites this paper.

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:56.299422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T00:54:39.349292Z digest=sha256:9d2731932038da50e4bb834b5bcd8b19a2568de77ceffce3511af7a91504717a

Observation b2396714-c593-4147-b06c-da2776362bf8 · inbound

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors cites this paper.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.271730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:9ac50ba5a5150bad56b3d9cf5cc16de59b38ed620a2d80ec86040a8ebee09265

Observation a6579577-05d4-4a44-a098-bf48c50f0bf6 · inbound

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR cites this paper.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.408581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:b80f14f08ec2f83ea744cdcb69f4f30fe8f33ddd27e4eac7bb759e49906c809a

Observation ed419b70-ffc5-4eae-bf39-c060302b0ee9 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:07:09.183624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:57:30.877915Z digest=sha256:39b0bc6eb49738a378399036fdb8a4abff131456fc93ef83c415c20baf784eff

Observation 4efd0409-3d61-4fff-90f7-0358f16c0b19 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.563663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:16b28e51fbca55bfe6d0076465d38fe662386745459726c67b67e044b9f87761

Observation 9a1a4a03-0132-4c2f-b4ba-3a3c836b1e46 · inbound

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs cites this paper.

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:44.104548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-20T20:09:42.841695Z digest=sha256:e67e4f8c56aab17ac4ff3d8ff5ea6a47a30f322e833b0c66ece196ced32dfa68

Observation f26dcbd3-d925-49b1-91fb-13e1c3be1936 · inbound

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning cites this paper.

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:33:04.068835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T05:30:37.685873Z digest=sha256:a20601709f0c116dd7a62daf4d22af53f2518f67cd8962c7010c906af081a613

Observation dfe95361-ed25-42cd-8246-efab17eaac41 · inbound

Finite-Time Regret Analysis of Retry-Aware Bandits cites this paper.

Finite-Time Regret Analysis of Retry-Aware Bandits Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:03:59.249766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T06:01:17.988127Z digest=sha256:f3354d0c94acd6354c88bddd07b391aab9d96d1f5c6ab57d647b585078648e16

Observation f35f1a59-919b-493e-b88b-69a2bfa9def1 · inbound

Finite-Time Regret Analysis of Retry-Aware Bandits cites this paper.

Finite-Time Regret Analysis of Retry-Aware Bandits Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.515573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T17:32:02.539675Z digest=sha256:22157ea08d74d1c432c92ae9ac840789f145721b9d9a3259c87dd218aeea41ce

Observation 42e13782-b429-419b-8dc1-0dce22bdd636 · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:41.116576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:ceb07b175c1023d78f984149a0df84279564bc4ccdf6a93d73a00e06e1aa5f7b

Observation 97047d31-dc3f-44fd-ad7c-ac45de92f9ca · inbound

Residual Skill Optimization for Text-to-SQL Ensembles cites this paper.

Residual Skill Optimization for Text-to-SQL Ensembles Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:41:17.254306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T08:38:41.126772Z digest=sha256:118c20cc34ed54beba9c8bc15ec49600a42a7b3155f9cb5febc83c541c3f26b7

Observation f655d136-4f42-443e-b1fa-a80c77dff8a4 · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.949295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:ee305eba2b9e2eea8d00c1d2403d2561952343c2c6bcb5d9f80aef59b93bf3de

Observation fc4b7227-bd3d-41aa-811f-d456642b1cf9 · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:27.306210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:d71ffb82be80c67e49d652adb89cdc631e504dcc5d964d2345488cdc2d2d87ca

Observation 19363c93-9a7d-400c-9e42-7115da93686e · inbound

Retry Policy Gradients in Continuous Action Spaces cites this paper.

Retry Policy Gradients in Continuous Action Spaces Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.024097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:550a8d147b525156e431118be83c35da8885613a48e87a514b0a21e52257492b

Observation b02f5bbd-79a9-4cf7-8857-b3f6bccb225e · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.383237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:ae21639cb47bd1a982fb76202a93e69dd1929fa19ca7a8741e68f43ae66d1362

Observation 742e7457-58dc-4252-90bd-f8328b9c11b5 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.976361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:eb7d596c99a839d6ccadc2fcc1e8e83f7f74c2da3752348e971457a5cb59432f

Observation eeba6203-dd94-4814-ae96-c7756f3bf42c · inbound

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR cites this paper.

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:55.621080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T01:15:18.237335Z digest=sha256:d962b55a74d02c4e547c1edad4893a69569eb4df081fa50e952bc5a6659456ce

Observation 31da5b42-3f07-437e-aab9-7e9d4d976486 · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:14.053289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:4a2223514ec0e3eb095a7085436fe41e96a6b1efb469b2922ad2da9baf21c898

Observation a43c264f-4835-47fe-934d-7d24b231d5d5 · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:16.717650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:9d2eac2ff57347bbf0e9eb3bb3da17d6dda5e7a0715f748629d900666045e965

Observation daa031ff-e73d-4e70-88eb-e64ef10d3383 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.459905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:320ffab7b85447cae4bfdc25939aab66327f38211eab291995684667afbbb321

Observation c3aac177-1822-4da3-9390-aeb872f3e1f0 · inbound

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL cites this paper.

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.650057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-03T20:59:57.539909Z digest=sha256:bbea9883ea7896700917ebd0f516aaeb8e4995856a72540c69bbcb6cb7cf8aac

Observation 65a40f4f-3a54-45ee-94f8-cb0002af75a6 · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.768877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:3c859544cd2efd566a8f0beaa6357e163fc671bedd35d4a9ade7d9a934b5aadb

Observation 735e13ff-2756-48e3-88b5-344b01ad6948 · inbound

Spectral Rewiring for Exploration, Purification, and Model Merging cites this paper.

Spectral Rewiring for Exploration, Purification, and Model Merging Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T05:08:55.438431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:08:55.438431Z digest=sha256:968b5ae8967acdf2ecc3b430790aab73b0fd33cfddf9bdf72732452de6643187

Observation 11a1f182-415f-45d8-a3ec-9b9c8e7a7bd0 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:36.599379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:36.599379Z digest=sha256:11a133dca49903360db8b501418c30df7de94d6363d32f8aaa66a89ab6399f38

Observation f3e1b247-a286-4779-8f1e-f44820c54e7c · inbound

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation cites this paper.

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T08:41:09.320896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:41:09.320896Z digest=sha256:1a282a715b999cb5006223bfadd8f0832196bb5dbc771862a5104d1fa940987e

Observation 53d17f62-7644-4900-8a92-e26913548d27 · inbound

TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation cites this paper.

TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:52.387697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:54:52.387697Z digest=sha256:97f9a9223fd558dfd8183e497e578e13328537fe69269ecb521b66863249ba26

Observation 1f503e5d-6900-4a9f-9df7-401fa10d48df · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.330187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.330187Z digest=sha256:7c508ad41a3857e3fc321800d594fa5e687ffcc84045fe0335dfe561ba4e1f23

Observation 9cc2c55c-b733-40bf-aafc-e98d3ce6213e · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:16.232644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:16.232644Z digest=sha256:5dea308c004feb1038293fcbb9b75ad02e35809b27b441e36de4e420ef2327c5

Observation 53186b75-65a6-4686-bc35-4da457bbb35e · inbound

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning cites this paper.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.174330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.174330Z digest=sha256:a7bb0fb3767b60e66ab1bd59eb809f4f8206b48e643b75b7a0d7078784c1911a