Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:29:01.355429Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2607.26253.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:29:01.355429Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f23f478a-906b-4dba-8968-13c4034c7211 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fec385-0f6c-4a7f-ad42-4db7fd3f3f69 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Defaults and history warm-start.Defaults: B=64, k=8, n0=2, τlow=0.45, commit_min=k, prior Beta(1,1) , safety budget cap 6Bk
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80cda614-70f7-4fb0-94b9-7af2c0d2baba · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a215fe56-9b28-4714-8304-5cbabf1ac35b · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR OpenAI o1 System Card
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20c8011-8bc2-4e66-9541-d47de759d950 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Let's Verify Step by Step
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a342d7-a398-48f5-8ca1-385ab2de3150 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42201871-fb30-4c5a-ad08-3988bd9e9c67 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR HybridFlow: A Flexible and Efficient RLHF Framework
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89732fa3-96a4-41fc-be18-bd4b341bf644 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50571580-36a6-424a-8913-78de2b074c5c · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6913b44-d6af-4024-8b77-0ff5d948533a · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Qwen2.5 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a948e6b-bf0d-41df-8312-9c2aaf43771d · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMO: Less is More for Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338eaa19-99eb-44a7-8929-46ed9e353858 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0721e090-647b-4ff0-8756-616091f58f48 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea09dec-71d9-4b2d-9af3-1544d52c4ef9 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f751cbeb-b77d-4a11-9dca-32146724e119 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86be6341-bbd2-4899-a9b5-1988a1810648 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65024ccb-bf7c-4201-81e3-831653382380 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR (ii) Rollout allocationsets how many rollouts each prompt or trajectory prefix receives (Zou et al.,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d05d1d9-2726-4d08-9d4c-4cedbba0ae2a · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR These all commit budgetbeforea group’s own rollouts are observed (evaluate-then-filter even pays full groups for the prompts it discards)
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d4b39cd-b87b-4933-b6a4-f9311b6897a8 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107fcae4-0df2-49d0-b34f-af1453ae0ee4 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR effective
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b913fd-6f47-4257-9f9b-58bbc8cb07c8 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 1945
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e536f69-a577-4dbf-aebd-b9ed1efaece0 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 1972
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe76079-caa6-419c-a7f0-f4b776a83209 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1979
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27429d3-9be6-4af5-bdb1-fb5f2928fe5e · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380,
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a5b7f07-c35e-4f2d-b4f6-3fd4177e587e · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Proximal Policy Optimization Algorithms
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766ae64a-0c69-4538-a432-6701f12ac801 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92f760c-8126-4a5a-9a47-8ad5b54f81bf · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3843b153-84fa-4fee-90af-fe3a27f6222c · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMR: Less is More for RL Scaling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 069b4b90-4905-48ac-95f9-942c72db07fd · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Cppo: Accelerating the training of group relative policy optimization-based reasoning models.arXiv preprint arXiv:2503.22342,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1edf779-0458-4b8a-8c0d-1383f6c1e6a8 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc8e48d2-c3bb-4426-b1cc-fa095e22f5fa · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dced1ec-316c-45e3-85e1-ccff69ffac64 · outbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR No LLM was used to generate experimental results, proofs, or claims; all theoretical statements and their proofs (App
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.