Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:41:30.443487Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 8 inbound Pith citation observations for arXiv:2505.21178.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:41:30.443487Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:54:17.656180Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:19:43.875028Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b81ee456-0d52-4b5f-a579-962eb784e605 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e31e88-5060-4b0a-8c26-a120743c7073 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning s1: Simple test-time scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1420f42f-046d-4d1e-b85f-f8012d519340 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Chi, Quoc V Le, and Denny Zhou
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b96ee8c-6b8d-4c54-86ce-3f52d171f55a · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning OpenAI o1 System Card
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d1973f-f508-46ff-85b6-373f6c750ef6 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f177744-3d28-4f1c-aef2-1fe4e8ed9d75 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456363e9-8d2a-40f7-a591-f83ee6693c8b · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Demystifying long chain-of-thought reasoning in llms, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c3e0a8-97dc-42e0-a5bf-afd2471842bb · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7f16d0e7-aa85-44dc-b555-33bb96a862db · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9fd87bd-9edf-4305-8e57-5b10c52f7099 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Gonzalez
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b103f33-1860-4f6d-b835-a1d09a3214f4 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5192dda7-3a17-4409-a2a3-95e4236bd339 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14d24b52-b044-43cb-b105-0d52f8507a96 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a511c874-76ae-4ea0-85d0-2e683ec10817 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Concise reasoning via reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268f1b66-93e5-483a-b060-83623e1f2a68 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78ba4d1-5f53-4b63-b769-b4ea1b998dac · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7c4d8d-a154-4858-b2e0-37ecac62f55f · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Solving quantitative reasoning problems with language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8570983-096a-48d2-96c9-cb11b5ddd572 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83cab6b-6cbc-47a6-89d8-dc83d447fa79 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed5ca11-5da9-4e14-ae6c-686c2603a8b2 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0be12f3a-3998-48ae-b6c5-b0585ec4b4a6 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ada641-80cc-400b-98b9-b340269b483c · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267993ab-5ae7-41f0-9343-3a739f73831e · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a9d0b3-0b5e-4d55-9cc0-2409c6bff85d · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f985c0f-0f15-43e3-9239-8d282d54aa74 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa24d0c-3be8-487b-8067-7af43a5e905e · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 21283f7c-89e6-4307-8f3e-8c67dc08414f · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 166f3aa0-d6bd-4d93-93f9-12be5dd032c0 · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a85ba105-22b8-4869-ad5a-b8688d9142cc · outbound
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788f827d-c11a-4f5e-911b-d40096de2272 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1a6a0e55-3b5f-4c3f-8143-3877a86ae2c4 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 165
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8750ce5d-5a39-4973-b88c-f5a6ab0bb8e8 · inbound
When Less is Enough: Efficient Inference via Collaborative Reasoning Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation db530fcd-f20b-452a-a13e-c6bf8b302f72 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 251
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e7e74023-cbfa-40f8-bf5f-ce48eef83c14 · inbound
Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 95d5fa2c-63b2-4dac-83bf-2b0e9d95bf91 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6967e08-c3c8-4e23-8297-c4dd9a0ee001 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3634661-fae2-4be1-8448-6bce7fdca490 · inbound
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.