Pith. sign in

Paper Citation Record · LEDGER

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

As of 17 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 30 inbound Pith citation observations for arXiv:2506.02177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02177 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:58.345391Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:19:46.719051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.849796Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9831c0c-05e7-4a31-a01c-0fbc8f226c97 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.422853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.422853Z digest=sha256:d9e84493315f9afe251ef91bc43b80a656854529bdc8f98d36f4cda494a0464f

Observation 0d8a338a-89e1-47db-ba73-d11448e44cda · outbound

This paper cites Training.Our method is implemented based on verl (Sheng et al.,.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Training.Our method is implemented based on verl (Sheng et al.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.133936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:57.817375Z digest=sha256:53be6f81630bd508c4e081ff210173549816246e768a2780aabc041a3d9fd5ba

Observation 23c8d49a-973f-49ee-95cc-884995a9bb2d · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.037842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.037842Z digest=sha256:5b0d3bb72485fde34722dfd94c2fd8386a84813fd6e3471773a80ce13c7bdc1c

Observation f9c3030d-ccd9-4bcc-bafd-d5082f7d4ebe · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.160410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.160410Z digest=sha256:3fb3083cfc3ed3783547c151337b232d1cc8247f7d7ec0c3a524492f0b289f76

Observation 1513ae04-df18-4551-b3b0-5004701717db · outbound

This paper cites Data-efficient finetuning using cross-task nearest neighbors.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Data-efficient finetuning using cross-task nearest neighbors

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:01.140721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:54.446718Z digest=sha256:f530f55e8e8a8471d70a6a1b8113046d179d756a7a8f3b6ebf2ad2b244e477be

Observation 9b82e5ce-5627-4432-acbb-de0cc0225403 · outbound

This paper cites Large-Scale Data Selection for Instruction Tuning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Large-Scale Data Selection for Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.560575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.560575Z digest=sha256:711c12975e42ab6b5e2f399967517ca5ea244048ac5597c959aba3aaa1ac8504

Observation aedd8de2-cdae-4679-b01c-fc73610148f5 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.702600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.702600Z digest=sha256:bf98eccb54cacfe2d59fc0f60109c4ba224b8f353f7bfe954cdd6fcb63b1bb0d

Observation 04d656e2-c62d-4612-92f6-7b1e770b01bd · outbound

This paper cites Let's Verify Step by Step.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.000411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.000411Z digest=sha256:1702633dfad0254a53e96a546e6d5526ca1284227c60604761f2a64a0d9e7b5c

Observation c30c30e7-a2a5-40f5-aa80-9eade33d7799 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.145743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.145743Z digest=sha256:4143fbbd8d774e2d4db7171c8b80b16b744d39bfaa58e1099e55d7cd937de0f9

Observation 0f96e0ce-7823-441d-bc85-09fc21a62170 · outbound

This paper cites Enabling Weak LLMs to Judge Response Reliability via Meta Ranking.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Enabling Weak LLMs to Judge Response Reliability via Meta Ranking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.280113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.280113Z digest=sha256:1ec03357508badf6da6d59053ba16f1e4c1de6fdff5bfb894a03d111fe0e05ee

Observation 281fdffb-7daf-4088-a06d-b603cad1e1f9 · outbound

This paper cites s1: Simple test-time scaling.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts s1: Simple test-time scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.416178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.416178Z digest=sha256:f6d53776027d1fa8837ffb6e5d81cddeb53674c215101a598b1f5d6a1ce42479

Observation 6d8283ee-dc53-42bc-bf0b-7e40a6c8fb88 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.521985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.521985Z digest=sha256:8615b91aaf2f2dccc22422881ac0d9d930546e9a7b30de20380a0314021ba2f1

Observation 8325e58a-b228-4603-8f90-fa9d1df8c754 · outbound

This paper cites OpenAI o1 System Card.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts OpenAI o1 System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.643161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.643161Z digest=sha256:195982083d2b9d457bc82054a85f88eb9cb12edbc47587f952f57fe396fa2cca

Observation 58fd25b9-d2f3-4476-98de-d4e90c16c08c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.813166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.813166Z digest=sha256:f5a25dc912d8e34f1c2e516f30dbafb8b05ce2814df12b8b16ddbabe0ee19a9c

Observation 10a3f3e9-9bb3-4dc5-aaa8-8ef6b83ed13e · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:55.929448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:55.929448Z digest=sha256:c5d9bf6f43b2c804823f6b608651914ddad06049b6f6e7eb282bcf152107b8b5

Observation 8bb4b8c2-2e6a-4fec-8a11-2d49ab3cd961 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.042626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.042626Z digest=sha256:fc6c31d8f797f89f637f91a8e123d42ed16762251956842482d0e34f6e8e21e0

Observation b7f103ad-20e3-4ede-b2fb-bba1a62b64ea · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.370913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.370913Z digest=sha256:eea82a03360aba92d590b303e224770ea7c290faa0ec02481cbf0bccf1f44553

Observation f678aa86-9893-4cd6-8d94-f0914b3a7036 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.534466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.534466Z digest=sha256:cf720fcc3cfcaadf712aeaa6d476c7f3e3fb4c756e8e634ca4e889e0b5c43762

Observation 87883627-e99d-41f5-b282-cc9822319cae · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.659506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.659506Z digest=sha256:7d390bd43663728d88fdbd55f967f57062adae3de78a76c7935043c5970c4c78

Observation 40fb4ab4-d45f-4e03-adc9-c5a196ab3651 · outbound

This paper cites LIMO: Less is More for Reasoning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMO: Less is More for Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.783335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.783335Z digest=sha256:bfd6c94aa31d4ffd44c1a0f0116ec8bfbc52b172819eb5346c3ed46c380460d1

Observation 209e468b-62bb-4f5a-b75b-ea1eb0702bd0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.907899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.907899Z digest=sha256:70a4c19252d1249411d9b01445b260b789afb8fa815ca0879b9e6f5f6e1956d6

Observation 1871f2f7-c535-4ef7-9dcc-a2ad22c02ed6 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.031251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.031251Z digest=sha256:a6defa0706ee0c50c6bbf9aabdf0f16d9c423c9897486530f7d7d82f60e52eb7

Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.154091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.154091Z digest=sha256:89f9865b804cdf3f88afc11779af180c64555d7db805d1870b8371031dfa885d

Observation af96af09-a3de-4845-a58d-785ada5276b4 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.305024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.305024Z digest=sha256:46dc9e6482cb2de025ee94a1cae65975391d243040a7ed0cadb75c161082b927

Observation 39da3963-895b-4fa0-9746-a8b686108307 · outbound

This paper cites Coverage-centric coreset selection for high pruning rates.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Coverage-centric coreset selection for high pruning rates

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.880737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:57.447663Z digest=sha256:20a83d4e3093ddfef3951b883ce09fdbed70871a9a36fbfdfd185fdbd9260cb7

Observation b4b7d512-6da2-4e00-a08b-ae0adce931fb · outbound

This paper cites LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LLNL-affiliated authors were supported under Contract DE-AC52-07NA27344 and supported by the LLNL-LDRD Program under Project Nos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.655878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:57.550465Z digest=sha256:3ee96b03ea275833ee323fd155d90e98709dd75f678c7732c94de8bc829fb7cc

Observation 9d4a53b0-6123-428b-ad5b-49517ad0fe3a · outbound

This paper cites We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We find that training on DAPO alone can degrade performance on LaTeX-based benchmarks, so we augment it with MATH to preserve formatting diversity and improve generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:34:00.370119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:57.673181Z digest=sha256:d1298568caa0a604397d3b9f7360b17aa07f64b0f2afada488f981136f047116

Observation 97e561c9-9c20-4a84-abea-f46fe94f1050 · outbound

This paper cites We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use 4xH100 for Qwen2.5-Math-1.5B training and 8xH100 for Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.873088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:57.964485Z digest=sha256:2cef6ad7a0b7bf81fbfe0c1b9d32d3a54648462b8d455d276ad6cf025429179a

Observation 7eafab92-670f-4b36-bd34-28395fb1569c · outbound

This paper cites an unresolved cited work.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:33:59.357792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:58.212664Z digest=sha256:7c49edb8ebddaa0fda8140c9ff08f35da561c505be6b571d00e5f360fe8b0547

Observation 64209bf2-bcdd-4765-acb2-35b1ea20ce8e · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.510451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.510451Z digest=sha256:a45fbcbf247356ef0634b5fc3fb1753d29834ae17accb210f102b73568fb31b9

Observation 8efec914-804d-416f-bbd1-df6fdc0ad840 · outbound

This paper cites We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We use β1 = 0.9, β2 = 0.999, and apply a weight decay of 0.01

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.602692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:58.080472Z digest=sha256:b24cc1315255d12f40d46a4cec119900fcc041e938171138f135ebce96ebac40

Observation 2a05e5e0-9c29-4444-89dd-426e927ce2b7 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.204875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.204875Z digest=sha256:8b98f85f7849491ff0dd64172fe6cfc07e82fff15e2f8efd02bf1fe3c8fbd17b

Observation 078a963e-511f-47f7-90a5-90629883b038 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.297876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.297876Z digest=sha256:0370ddfd417644fe7e7216a69a4534c85a5674a9230d0ef858a9a1d4f889f782

Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.876617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.876617Z digest=sha256:36e7b0edeac82ad1956f523af71a0e16053e431aee3a6b9d98456efffce0e12c

Observation 5182579f-4527-4fea-820f-fafc04ade777 · outbound

This paper cites Active Preference Optimization for Sample Efficient RLHF.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Active Preference Optimization for Sample Efficient RLHF

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.628034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.628034Z digest=sha256:4045f989e515d267c506b20add4b9efa533ec6f808bbbf5ced98b6f36860fb1d

Observation de08b487-e30d-48b1-90f0-2e748fca46e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.786117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.786117Z digest=sha256:ef5788740a241e3c8529eda9d2d3d0ff753ece55215446d9865d7824d07f9208

Observation d5eb21f6-e440-4230-a022-5f7f5ebe07d1 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts Open r1: A fully open reproduction of deepseek-r1, January 2025.https://github.com/huggingface/open-r1

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:53.916913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:53.916913Z digest=sha256:91364397dec0fbe223d8db7ef7b44f9ff94655cfc948723d03eb1ef31b95abd3

Observation 9d935f1b-5438-49e7-a02b-8e5183d840bb · outbound

This paper cites We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts We evaluate all models with temperature = 1 and repeat the test set 4 times for evaluation stability, i.e.,pass@1(avg@4), for all benchmarks

Reference 8196

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.131928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T11:33:58.345391Z digest=sha256:23c26ff2337b980da0e7fb82c9b474bf59238288acecbc0f74f9a1752c1a04ac

Pith citing papers

Observation d084dc3c-910a-4cbc-9535-8d8786694e82 · inbound

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance cites this paper.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.719051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.719051Z digest=sha256:e6f55a650fa223937c991ad0d78e0669c0800f3fcc40578981e48f3178a6c880

Observation b4d40336-1471-4269-b24f-77da40690d0f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.638499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a7c452b277586f8e9759073b0caafefb550b74332108e3432878f4fb28d2a468

Observation db8baea6-a880-4384-a29e-9fd1060a96ae · inbound

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO cites this paper.

LASER: A Data-Centric Method for Low-Cost and Efficient SQL Rewriting based on SQL-GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:02.195880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:15:04.225059Z digest=sha256:3f3d152a8ae854ce8338132dc7dcd2a20353cb66f7a7aacded2a4f197e4f6f96

Observation 332989d2-bc58-4d9e-b58b-55c974b8e2be · inbound

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models cites this paper.

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.840498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T07:05:45.081955Z digest=sha256:8f4168456818e97efc8b9f5df2769b6b4ef48848beed9f23818c92181f291858

Observation 3644689d-62f7-4cea-8a5c-5ec10e9c720c · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:36:30.855968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:b7112d2b9e746910a44413393d52a3740ad4ce43572a1660c0ebe524ef8ffaca

Observation 920e3085-d9e2-44c6-a90f-4243e10ddbe6 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 184

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.463309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:4402e619a17be3bf0f48df3ad82f585c980464b347fd3784089820462531fc0d

Observation cb3b60b6-8cee-4536-9fa1-7a060ba2fb04 · inbound

Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL cites this paper.

Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:09.059414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T15:31:23.161009Z digest=sha256:468ac90cfdb5bc7f980ab4035b474b35ad0ea51ee3519fd1b14c57b76735dfd9

Observation 1bd91141-483c-4119-97cc-f75940b57616 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:15.988767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:a9b619a358085cf5093aab76432e1bb4a6c6de8ecf07d90b0a075fb82edd9cac

Observation 99e4c427-1350-4aaa-bef3-db35fa1bc34b · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.200705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:ee3d3e52188aeb550e17bcd9568c24c821a6e3e30ba82cb4d7738b3404b2c27d

Observation e8dfdb1f-37e0-4527-abcc-0816e3b30fcd · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.733368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:e0bee5ff7caf49867dca0fd3e67d11cb8649145b5971f0654d3f42e6308befca

Observation 56fe42d0-5d57-4e03-8d3e-2fa1432befd8 · inbound

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States cites this paper.

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.882249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:08:28.113371Z digest=sha256:a2d6548c26d32992c7899d00d4d1f703b54891af00fc6b114a3c60c520cbf18f

Observation 1bb2620f-65b2-4965-a974-ffd22c58c6b1 · inbound

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States cites this paper.

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:45.835310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:01:25.989133Z digest=sha256:35e3a86a2600b02ee0401b6fb308283bccf290a82101ef1dfef7f7204663908a

Observation f86cb607-eed8-4ff2-a252-14fb94b01f17 · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-05-12T07:46:28.647091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:a8851a6498addf807132a0980bff471800bef5ee7aff3711b6192a0ca245e96a

Observation 12169cac-4b43-40bd-ae4d-62099273376f · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.327965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:3c61204605a2ab60d0f70530134aa94386a23491b726c8edb53e59a032256e93

Observation 117727e2-7d6b-4893-9a70-1b6ea6b10b7a · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.806520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:51c30eddc93f4ac5adebfad3eb8ce65742e30854d128027cfb260aa0b441a88e

Observation 39a0d34e-3e63-4bc3-b2c7-ecb5096dfb5c · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:14:52.595524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:63176eaafe8fe6465c12b352749e3fe4f619f3e28dae5cda3ad9b6ec3d166f1b

Observation 2bd94e0d-39a0-4977-a38e-91ea45ade716 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:39:00.348750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:0bf85cba4de779c248dc582bb3daab93b08b5b67c7c2a6fc20f3c44ade24f644

Observation 72467af6-8d03-4f29-8f8f-1424b2767578 · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.788110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:c6c9f0b924960d377b9bf0f8a5516ecdbc3a574208eefc586f4bbe4fcfa16cae

Observation 942d9897-902c-48e5-9eb4-31e12c4ae2cd · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.782288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:7ef963a84229f0195ca757981714a31b02d36fdb89bf6e62491c151fc69e953a

Observation 24f670bc-ea5d-4b7e-a474-bd4cdf87ed2c · inbound

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training cites this paper.

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:15.361248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T08:38:52.411671Z digest=sha256:ad5128444845696c7863290a1a280fac166cc70d0ad8e628add4d027d6725575

Observation ba73a442-9db2-4e22-ab0a-1a24cbdae46a · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 70

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T21:16:13.345991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:b2b06f64890bc6b97db9d7af8e34c53bf66ab6cec86f9cc2cfca190e1212e219

Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.403920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:afd171c4c65a6b8802c85ef49e460edd29ab7a129308b4179720e141574010f2

Observation 9fe1ca1e-f844-438f-8466-60f0a59ecddc · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.917602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:695e3b1ea3dadd6990b33f006afd99dc09c89a45b44a01366c39c87da86a5dcc

Observation 342d495b-9320-40ba-adb9-58a2db117d97 · inbound

CATPO: Critique-Augmented Tree Policy Optimization cites this paper.

CATPO: Critique-Augmented Tree Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:47:28.013303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T19:30:07.971424Z digest=sha256:a8b0be20d67e44f620e75eda8b4743c3bde832709b0f1260a6ded9fde59d319e

Observation 125b38d1-05be-4ade-bf2d-03ff4f6c5ced · inbound

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization cites this paper.

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:37.386022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T14:02:05.833651Z digest=sha256:72963f4ab3b8d580c4b901f51d150f22dfd3ba292d7a859f6a02036db2418e8c

Observation 1e5be0ca-5cbb-4df1-9a9e-4cb4ed65cca2 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.851250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:073b9b05ed7fb7224b5a87746b3dfcdd46d291941806dd78f0d6b22f1d2a59df

Observation 3ea09dec-71d9-4b2d-9af3-1544d52c4ef9 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.908981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.908981Z digest=sha256:9d5fb5e38928c0e238af050d44add086d48fc95d4c6399f38fd6c3eccceefc87

Observation 90c868fa-b460-46cc-b1cf-3190f1bd4b4e · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:13.934392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:13.934392Z digest=sha256:b0d0b4f6fd21a5ee17711baedbfa43ae214aebd23a115abf8fffda44a8aee750

Observation 79f0a914-cc9b-46b6-bf4d-f8132541aa86 · inbound

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance cites this paper.

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:11.829438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:22:11.829438Z digest=sha256:a5e3c1ce54d6c6002e61f661ffa684ff234f9e8842b748b600699e2920fe0b08

Observation 2ac7fa79-d18f-4e95-83b3-d54f63cf1327 · inbound

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR cites this paper.

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:22:55.603592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:22:55.603592Z digest=sha256:0d2169d0093955559f26474f8bdcaacb940aba05e456c1dc5a53649b7914043b