Pith. sign in

Paper Citation Record · LEDGER

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2607.26253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26253 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:29:01.355429Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f23f478a-906b-4dba-8968-13c4034c7211 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:58.772850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:58.772850Z digest=sha256:29a4ea2bdca284a19d3d2dd30bbcf2cb520ade40e0c31042465ceab2bca0b4cb

Observation 42fec385-0f6c-4a7f-ad42-4db7fd3f3f69 · outbound

This paper cites Defaults and history warm-start.Defaults: B=64, k=8, n0=2, τlow=0.45, commit_min=k, prior Beta(1,1) , safety budget cap 6Bk.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Defaults and history warm-start.Defaults: B=64, k=8, n0=2, τlow=0.45, commit_min=k, prior Beta(1,1) , safety budget cap 6Bk

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.283631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.283631Z digest=sha256:43f5be97c39b0b4f73b60b5b16e2975079498bb18390d1a35242cdc1edc39b6a

Observation 80cda614-70f7-4fb0-94b9-7af2c0d2baba · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.222616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.222616Z digest=sha256:22b1203c358f5d45b4160dc8a9e3aec9146486063a6603402a4cb3a3fa838941

Observation a215fe56-9b28-4714-8304-5cbabf1ac35b · outbound

This paper cites OpenAI o1 System Card.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.355379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.355379Z digest=sha256:d779b6f7aca25960057bb84ec049019a02ec87f8757064664adf98360bece972

Observation d20c8011-8bc2-4e66-9541-d47de759d950 · outbound

This paper cites Let's Verify Step by Step.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Let's Verify Step by Step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.555279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.555279Z digest=sha256:f7d22e82512e70d1eedb9b1cb767e32fabc39fa5d355f9cf242dbc3bf2372830

Observation 03a342d7-a398-48f5-8ca1-385ab2de3150 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.796151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.796151Z digest=sha256:7d6d44eae663c109a4df9c4ec02765d6db0625357a92ef47fd0ac63fc96bafd6

Observation 42201871-fb30-4c5a-ad08-3988bd9e9c67 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR HybridFlow: A Flexible and Efficient RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.198248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.198248Z digest=sha256:d7b7e15f07659314a687e8e024597bcefea74bdfd9c9434d360b5728a6b3786b

Observation 89732fa3-96a4-41fc-be18-bd4b341bf644 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.283036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.283036Z digest=sha256:a123ed639c7a2304b969409311e3a6bc145d22b2f016266ae0e0714f439b248d

Observation 50571580-36a6-424a-8913-78de2b074c5c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.379265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.379265Z digest=sha256:990806e40e3a5356549a8648381cff971c587f07765cfd1d67b4c2db45851b99

Observation b6913b44-d6af-4024-8b77-0ff5d948533a · outbound

This paper cites Qwen2.5 Technical Report.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Qwen2.5 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.565728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.565728Z digest=sha256:41eae00fea0f35f507deec4f72dc10e1bcd8a31685d858a7b4e0658cc3e87235

Observation 6a948e6b-bf0d-41df-8312-9c2aaf43771d · outbound

This paper cites LIMO: Less is More for Reasoning.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMO: Less is More for Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.654824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.654824Z digest=sha256:c5d1e0f610760fe563659112a5555273e9550069a5b778a6247253788d24ffd3

Observation 338eaa19-99eb-44a7-8929-46ed9e353858 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.753300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.753300Z digest=sha256:69f1b59557695a96061c11d06dedfb1074edf2e7f7da20cc45509d3c5e5e2fd2

Observation 0721e090-647b-4ff0-8756-616091f58f48 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.841014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.841014Z digest=sha256:3ecdc1c3babd3bcc1484f8cc29c08384237147b0b19d8ae3676bd8e9ca4fafa6

Observation 3ea09dec-71d9-4b2d-9af3-1544d52c4ef9 · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.908981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.908981Z digest=sha256:74bdb6996c66c28ba0237e95a61531899f839031a21949e11583b06023913994

Observation f751cbeb-b77d-4a11-9dca-32146724e119 · outbound

This paper cites TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.970654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.970654Z digest=sha256:0326284abbf005e52f4083d41f062a017d27da50c4a2afb54850d466f6baa947

Observation 86be6341-bbd2-4899-a9b5-1988a1810648 · outbound

This paper cites an unresolved cited work.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.076408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.076408Z digest=sha256:990510df8e092350da7b16313c8c981e1c2609ef37b879fd1f3643d91c662233

Observation 65024ccb-bf7c-4201-81e3-831653382380 · outbound

This paper cites (ii) Rollout allocationsets how many rollouts each prompt or trajectory prefix receives (Zou et al.,.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR (ii) Rollout allocationsets how many rollouts each prompt or trajectory prefix receives (Zou et al.,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.128562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.128562Z digest=sha256:a90f8ac4eb9fb8597f1c536325f85cce4955e20e6e246c2b43f3e2b07cf6b291

Observation 5d05d1d9-2726-4d08-9d4c-4cedbba0ae2a · outbound

This paper cites These all commit budgetbeforea group’s own rollouts are observed (evaluate-then-filter even pays full groups for the prompts it discards).

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR These all commit budgetbeforea group’s own rollouts are observed (evaluate-then-filter even pays full groups for the prompts it discards)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.176326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.176326Z digest=sha256:5f00f2215ab7f9119e29d7d6c5b66bd1db223b11453b4f4305e8f5e152d3c15e

Observation 9d4b39cd-b87b-4933-b6a4-f9311b6897a8 · outbound

This paper cites an unresolved cited work.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.227004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.227004Z digest=sha256:67d5ec6b971f72054e1b68ed3f01dab5bbfbd532b387c9ee1dc06666ffaca594

Observation 107fcae4-0df2-49d0-b34f-af1453ae0ee4 · outbound

This paper cites effective.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR effective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.355429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.355429Z digest=sha256:757961f48fec1254208c3a7d31911605fbc32da1a52ecdf55631ae10d90bf011

Observation d9b913fd-6f47-4257-9f9b-58bbc8cb07c8 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 1945

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.474113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.474113Z digest=sha256:fd4fae46c13dc1525ab9a16bfcc1241d6260247d21814de82dc71a69ce55389a

Observation 0e536f69-a577-4dbf-aebd-b9ed1efaece0 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 1972

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:58.867107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:58.867107Z digest=sha256:5d2f9f5d78caa2a048c1cfaa756fbfedf557cacb5fab34bef7410d13e74b2671

Observation bfe76079-caa6-419c-a7f0-f4b776a83209 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1979

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:58.901148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:58.901148Z digest=sha256:ee3dd48a335ffc455b13a1511f43d69906046a6c08e0dd80c4a94e093e10dd6c

Observation d27429d3-9be6-4af5-bdb1-fb5f2928fe5e · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380,.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380,

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:58.813150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:58.813150Z digest=sha256:e436db297af3477678c062238bff6f02f34236d89241504bad412e99dff21628

Observation 4a5b7f07-c35e-4f2d-b4f6-3fd4177e587e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Proximal Policy Optimization Algorithms

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.925404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.925404Z digest=sha256:01cc04944c1eaeae17e24214c51086416f5c3d37349d56367a161797f571bdc2

Observation 766ae64a-0c69-4538-a432-6701f12ac801 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.088737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.088737Z digest=sha256:dc7d43b2373059fd4f1e4b22f3d88004bbaf3957ef0784736cb2a3ffc71d536f

Observation b92f760c-8126-4a5a-9a47-8ad5b54f81bf · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.149242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.149242Z digest=sha256:9043d59d8f60408ce6c512bbc51046b4c418043b11eec8dfd310058237ad52e1

Observation 3843b153-84fa-4fee-90af-fe3a27f6222c · outbound

This paper cites LIMR: Less is More for RL Scaling.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.453562Z digest=sha256:33894b1841cebe6d59584f239b9e9fdd67e8c2c246a8054310a93795d8a99527

Observation 069b4b90-4905-48ac-95f9-942c72db07fd · outbound

This paper cites Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342,.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.637477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.637477Z digest=sha256:5198dac4e37767f576501f6a52ca4921f50348e86b01388410ae8b7404dd796d

Observation a1edf779-0458-4b8a-8c0d-1383f6c1e6a8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.008112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.008112Z digest=sha256:3dec6004be66861846178cd80b30588a339069ebe4edabba41321756d8722add

Observation fc8e48d2-c3bb-4426-b1cc-fa095e22f5fa · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:58.939296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:58.939296Z digest=sha256:22171f42f589dcb856081960ae4b5945a2979e20549b31e193562514ea1726c2

Observation 0dced1ec-316c-45e3-85e1-ccff69ffac64 · outbound

This paper cites No LLM was used to generate experimental results, proofs, or claims; all theoretical statements and their proofs (App.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR No LLM was used to generate experimental results, proofs, or claims; all theoretical statements and their proofs (App

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:01.024615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:01.024615Z digest=sha256:9373d939e718ecfdfe5f32ce698fc2c8c05b5940202317d39376018c7879b985

Pith citing papers

No inbound Pith citation observations are available.