Pith. sign in

Paper Citation Record · LEDGER

LIMR: Less is More for RL Scaling

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2502.11886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11886 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.318585Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.759210Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e60bd50d-4186-4064-8860-005c7aed1ee1 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models LIMR: Less is More for RL Scaling

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.483767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:7558541919b01e9a9eae618982f037a5fdb2649731436ab82a3666db711519fb

Observation df54a10c-c8ad-4ce8-b3ed-9ba41dd7c994 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering LIMR: Less is More for RL Scaling

Reference 180

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.318585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.318585Z digest=sha256:c59ed8c38a27ed81883d200d38adbcdb92a681e9ffd18fad0e0b87248c3c430d

Observation 8653614e-8e49-4678-92cc-08bf6d9a94ad · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs LIMR: Less is More for RL Scaling

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:26:48.445619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:e88692b81280684bac35eec21a290c3804c6609f1105ac64d6b0eb2e3755cf4c

Observation bb5e8f2f-9fcf-4fb2-a0e4-46bbdeb4f310 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models LIMR: Less is More for RL Scaling

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.239798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.239798Z digest=sha256:176450a3a96dcd98fe022b493b00ea70d3d651f0c9289fc392826ce10786d583

Observation 0c23b6d1-0c1d-4881-b783-53a571ddf733 · inbound

Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey cites this paper.

Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey LIMR: Less is More for RL Scaling

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:02.097662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:02.097662Z digest=sha256:6f8b4e7db7baa6661f50fd3637daa5136580ce8e0f3720fddf02f694297d8b52

Observation 3f5f2c92-a122-4181-89bd-cb394fe29eda · inbound

CEC-Zero: Chinese Error Correction Solution Based on LLM cites this paper.

CEC-Zero: Chinese Error Correction Solution Based on LLM LIMR: Less is More for RL Scaling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:44:05.730793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:44:05.730793Z digest=sha256:75fbcf4311bb4c296ccf9c7c1c3615b00bc06c0eb88d016be5a2ee2132be170f

Observation 94e76ad0-6676-4397-b153-21dd6224dde3 · inbound

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning cites this paper.

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:19.835682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:19.835682Z digest=sha256:342793f52eec9100d93e8223238f1000aed5baa79912e20c5c0bb501c43f486e

Observation 63cfa6f5-6848-4bea-8218-27a564635c80 · inbound

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection cites this paper.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.207868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.207868Z digest=sha256:bebbdea001ac501769371d5de3b9fa99118b57092d6ca9cc6325093ddef9c94a

Observation 9418ddc7-8284-4590-81cc-d8bc205438cb · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.385679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.385679Z digest=sha256:e2df519b2b97b530db8042bf602b0ace8c2e39fc14dd8df1bf25121c4a92c7c7

Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.361350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.361350Z digest=sha256:03209a83caf6c99dddccb2c14ba88fd7d1f9ac6bf19d6c487eb86135167a1a8c

Observation 6f9b73f4-d090-4ab2-93a1-d8eea78a0da2 · inbound

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning cites this paper.

Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:41.222191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:41:41.222191Z digest=sha256:b557aa945a44a7abed62a4862ae34c1d86ca4c0740b214c4270bbb71da2e8b0e

Observation 12b3dd73-f9b7-4c0c-abc1-d33bcdacbf4e · inbound

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning cites this paper.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMR: Less is More for RL Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.358276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.358276Z digest=sha256:a52b73b0f786e36a5b4f95f2c1b7d3ea20172ac0fdd163036c99a8a0058c74a0

Observation d3903c7a-4f40-4b10-bb8d-1917083a9ea1 · inbound

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning cites this paper.

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning LIMR: Less is More for RL Scaling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:32.954737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:32.954737Z digest=sha256:87237881ed4e0f962ac10936d4258ad3f708da95b9fc03fc1adf00c32e86302c

Observation ba80164b-d1a3-4ad6-b274-f52d49fe5572 · inbound

HardTests: Synthesizing High-Quality Test Cases for LLM Coding cites this paper.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.430228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.430228Z digest=sha256:84b0e6d605804dfe3941669e74706a289a60ec2312313444af81a52e65ccaf20

Observation 622b3240-6a49-4b5d-ba52-4a2c48cb0687 · inbound

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs cites this paper.

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs LIMR: Less is More for RL Scaling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:25.160677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:25.160677Z digest=sha256:5c24b3f16daca6b641eec219f9c54bc35116b8e6fe1fcb9a1e05c257cbe938ad

Observation 07818c08-b6e3-4a43-a672-edcb1c042834 · inbound

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis cites this paper.

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis LIMR: Less is More for RL Scaling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.327280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:35:45.327280Z digest=sha256:68e7448b18e9a0cbf1753df3c32f53f82d02709a87da101b896bb4637665b7a4

Observation 5c790190-4f4c-4b4f-a4a2-43d7f28d5c46 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.876617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.876617Z digest=sha256:32c63fc7b49d82759f41d6509ab63f048d430c2b192241d0c877ed478b8666a5

Observation 86b8cbe4-6afc-4a59-89fe-b898730394fb · inbound

How Far Are We from Optimal Reasoning Efficiency? cites this paper.

How Far Are We from Optimal Reasoning Efficiency? LIMR: Less is More for RL Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.467346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.467346Z digest=sha256:ee7a191621900783bca1f4257bb0369e6c6230ba1c54fba327e818db00e765ae

Observation 67b2ffa4-f2eb-4445-bf66-67a400742f9d · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning LIMR: Less is More for RL Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.145061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.145061Z digest=sha256:2f9dc5373295bb5ad3646c004ea26ba11a13bd066b1ba393dddd2a267fa3d484

Observation 960841b0-a8eb-43b4-ae90-d6e7ef6b0273 · inbound

Test-Time Scaling with Reflective Generative Model cites this paper.

Test-Time Scaling with Reflective Generative Model LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.775814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.775814Z digest=sha256:7dd0fc490d5dbd0b01a8cd849a120b856e3f7d8b1ec150bbe5cd7a580be092db

Observation 0bd2e522-1f8c-4017-9ab6-ba400e554300 · inbound

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization cites this paper.

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization LIMR: Less is More for RL Scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:12.582797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:12.582797Z digest=sha256:237e35581ee39f0bca919a41d070868ecfdd5022ecd8598b6a2db8b706ae2ff2

Observation b894d50c-c028-4d0c-9c0d-fb4703238dab · inbound

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization cites this paper.

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization LIMR: Less is More for RL Scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:36.105216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:07:36.105216Z digest=sha256:820a66823d6d6f3dd6e373f9ee5142daba88872f9a75d056eeb3c844bdeebd4c

Observation fd345c14-e1f0-4cdf-9d3b-80b3094c52e1 · inbound

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization cites this paper.

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization LIMR: Less is More for RL Scaling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:31:53.188452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T22:26:52.748349Z digest=sha256:f6ed14dbc556bbd0c6a480d849f2ed43502071e45fe1b0aa834a0dd066827a02

Observation 9b49024a-7409-4cfe-968d-bb5055889cf4 · inbound

FormaRL: Enhancing Autoformalization with no Labeled Data cites this paper.

FormaRL: Enhancing Autoformalization with no Labeled Data LIMR: Less is More for RL Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:11:23.150929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:11:23.150929Z digest=sha256:fa559a93b89e390b2bef8b03c985bfe34242d5d490c75724760a36e215cf5853

Observation b7d6496c-3f15-4a00-b1e7-a273af4f0373 · inbound

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning cites this paper.

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning LIMR: Less is More for RL Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:56:02.527244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:56:02.527244Z digest=sha256:9e4500ea5bdca321ffc0750365ad09266ed9d0e0419d07b02fc9613f8e9d994a

Observation 8dcc5459-bfb3-4a16-b94e-a0db283a9778 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models LIMR: Less is More for RL Scaling

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.776325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:313c5ee474f921128b7fd6e04846a271d8669a000d3f5a31bb7cffa3b906bfef

Observation 3560fc99-bac1-426f-a0dc-af3069e0e8f4 · inbound

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts cites this paper.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts LIMR: Less is More for RL Scaling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.397811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.397811Z digest=sha256:2ba50ebaf8653fc7149666011e2c7694f7cab3b10a14130248f5ca021a8d2e61

Observation aeaa8e7a-bb26-4076-88c4-0d3f498b73b9 · inbound

ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation cites this paper.

ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation LIMR: Less is More for RL Scaling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:10:32.307963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T00:08:28.730095Z digest=sha256:634e10a5c78072ee15a8b55d1a5568d96324b85b304de7a2edf42e31bbd89f0e

Observation 79bc85bb-ab27-4304-95fe-6380e188238e · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models LIMR: Less is More for RL Scaling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.194777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:e87f079c575edb568b50c1ba7a24516625616fe1f1faf7663ada40f5ddcbb1e5

Observation b585bddf-8bd8-4bdc-bbcd-5039a2452c41 · inbound

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing cites this paper.

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing LIMR: Less is More for RL Scaling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:27:36.744086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T08:24:50.793194Z digest=sha256:c667b261d9869956e3815b92bb146c06f4c256ff4cf154db32634fc8f69d71d2

Observation a9a9ae24-6dca-4b71-9d82-5ae649740511 · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? LIMR: Less is More for RL Scaling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.331952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:79fd7affc26cd6b917fbf742e6a32ca64e7db1e944e1ec32f5240bc58d66e043

Observation 9f9d1155-8352-4010-98af-7708be565f98 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment LIMR: Less is More for RL Scaling

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.517501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:6cf463dfa4de9b4709cf31ce8919579ad526456cc87f839d45bc9ccfc3301f78

Observation 11c49c28-d055-4fb3-83ec-ea8d36bce8ba · inbound

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes cites this paper.

Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes LIMR: Less is More for RL Scaling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:10:09.661693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T04:55:18.468593Z digest=sha256:8476e170e2cae78862fe8f31bc0373f6cceb165922687331315308cd7cc57e8e

Observation 1509c011-042e-4c26-97d2-6a9e59acfd0a · inbound

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation cites this paper.

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation LIMR: Less is More for RL Scaling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:15.529008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T04:37:41.629935Z digest=sha256:8709efdc08a2f2e1cb1b9611d7b9b70b12344ddc5f87c8952a542ef7119636c3

Observation 27709112-a375-4f60-bd93-a7fc88a51927 · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning LIMR: Less is More for RL Scaling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.935572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:1487d2069f4b4d884befc7319a4a39fae07b106006e1b4d241f21a6d02bcb991

Observation e3478c03-e531-46c3-bb5a-fdfe21ce7e40 · inbound

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning cites this paper.

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:16:54.879375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T07:15:28.648869Z digest=sha256:307ec06187ef2324e7e7b660a6854ddad51370952741c939baf8aa928d205889

Observation 3f4abc9e-1b48-467f-8ca3-87fd222fcaf4 · inbound

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning cites this paper.

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:26:14.089922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:26:14.089922Z digest=sha256:82d6069a487ccb6f53dab89e4f9a3feaffc5989014e4c10fe43a935c927f1abb

Observation 61c068fe-f799-4ba1-a511-d6928464fd8e · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization LIMR: Less is More for RL Scaling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.798009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:04c041f503cb4ceba61cf994d0f0d2fdd38424ecd83bc768b36c9483948ce544

Observation 91d7a693-0bc5-4807-b903-a3f5c693b9df · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards LIMR: Less is More for RL Scaling

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:30.076748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:2d4e39f133edea10959c703be758c4ff3a89fedad7c361cc808a04b506549332

Observation d9a08857-37e8-4d7c-bba7-23faace8eda1 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation LIMR: Less is More for RL Scaling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.290990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:931ee2b0939583af86d49b024352200fcbf4995f2514efd767ec586dbbbb2d24

Observation fed1f1c4-eb3c-429c-a80c-6f075a26b65a · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance LIMR: Less is More for RL Scaling

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.101924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:2fcaf9207578d75030a0b526a3c1219a2008be111043888726705d9e353093ec

Observation e06bc65f-d8c7-46c4-9278-272df101a202 · inbound

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection cites this paper.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection LIMR: Less is More for RL Scaling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:13:30.088668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T14:08:40.968105Z digest=sha256:1dd9221ebbb9d7c31a6e893af7a8d67671b0de9500783065e1ecd95a9c277cd8

Observation e27e7df0-802f-4eea-a729-f5b9523d4823 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation LIMR: Less is More for RL Scaling

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.626999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:d8d2f23ab89e36307dca7d18aafac2ca15db81f5a367747e55dda49a2dc43d97

Observation 83f0ada9-65b0-4860-b6ed-92bbdbe41d9c · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:36.760518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:f226ded25c85dd73448f37534fad5bcaa9cbf12cdd1a70206a4a446b181913b0

Observation 89a16f38-dd58-40f5-96c6-ac16c8f5eeeb · inbound

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning cites this paper.

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:14:27.040572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:14:27.040572Z digest=sha256:64fde42150a823b6511620ac30c1c64cc89ec86d0d062af9266cc88e8bfa1a7b

Observation 3843b153-84fa-4fee-90af-fe3a27f6222c · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR LIMR: Less is More for RL Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:59.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:59.453562Z digest=sha256:3f0816a433039703b389daffa4c48b7b170d81745d0afbb322f16e401f812bf4

Observation fd9d861c-ef6a-4fd8-b5e1-be805572c6a1 · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning LIMR: Less is More for RL Scaling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:11.009307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:11.009307Z digest=sha256:e658f9b186a34ee335029b1b0187aeda6e9a780b8d32ac5b39c6339fedd120e4

Observation 7ac8c575-74f0-4695-83a4-4668ce82174b · inbound

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) cites this paper.

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) LIMR: Less is More for RL Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T01:07:16.546910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:07:16.546910Z digest=sha256:6cebd22c2f87e84ec1e7f5f5617138100f4f46bb5825c3b9bed115e6f3aaa34b

Observation a8024bfb-8703-4df7-be90-c56f4140dd78 · inbound

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility cites this paper.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility LIMR: Less is More for RL Scaling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T14:49:03.884561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:49:03.884561Z digest=sha256:3f718b835b421fe17a5b87d923d103a6e46f66f6792f75a91ff37c2594902514

Observation 2019c7f7-8f1d-424a-960b-4b1ca17e1729 · inbound

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training cites this paper.

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training LIMR: Less is More for RL Scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:30:37.394264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:30:37.394264Z digest=sha256:b09a766679af1a129d120121e5437662c97341e08bc4d57e5363f6a4318b7503