Pith. sign in

Paper Citation Record · LEDGER

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 8 inbound Pith citation observations for arXiv:2505.21178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21178 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:41:30.443487Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:54:17.656180Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.875028Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b81ee456-0d52-4b5f-a579-962eb784e605 · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.371918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.371918Z digest=sha256:337fe6d750489a7e09c531582f844ced2335ddeb4deb4aa02c68b9c9ae01675e

Observation 12e31e88-5060-4b0a-8c26-a120743c7073 · outbound

This paper cites s1: Simple test-time scaling.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning s1: Simple test-time scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.479540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.479540Z digest=sha256:518031e01d2220d387761e0e739a65590e282e29da7862d543bfd5149c504c95

Observation 1420f42f-046d-4d1e-b85f-f8012d519340 · outbound

This paper cites Chi, Quoc V Le, and Denny Zhou.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Chi, Quoc V Le, and Denny Zhou

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.585725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.585725Z digest=sha256:425eca1996a15d87eada1c2ae0f4996688fe37148df342c4b5312e7a3a31c36f

Observation 5b96ee8c-6b8d-4c54-86ce-3f52d171f55a · outbound

This paper cites OpenAI o1 System Card.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning OpenAI o1 System Card

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.688241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.688241Z digest=sha256:05cb9e0d0f3d4c448798704db905758540c7a7913392b07f3288b03e7e576c37

Observation 33d1973f-f508-46ff-85b6-373f6c750ef6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.743661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.743661Z digest=sha256:7a271e49a9cc4a452f630c227b4a97ec7346ad7d6a8029a64787af1acc3353c3

Observation 1f177744-3d28-4f1c-aef2-1fe4e8ed9d75 · outbound

This paper cites an unresolved cited work.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.831623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.831623Z digest=sha256:7d586c865761a4eb6ba6427f1e4aeb641830be1b090066570fc1b8b0e81423c8

Observation 456363e9-8d2a-40f7-a591-f83ee6693c8b · outbound

This paper cites Demystifying long chain-of-thought reasoning in llms, 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Demystifying long chain-of-thought reasoning in llms, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:28.905630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:28.905630Z digest=sha256:aebf09e4f527e92ac2c68cf034590963e421079738de56768c31aa4f708f8543

Observation 86c3e0a8-97dc-42e0-a5bf-afd2471842bb · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:41:31.410685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:41:28.965945Z digest=sha256:58f171a3bd4d66ee0c842757332b316aa202566005e7706fa3967dc6b9238b9a

Observation 7f16d0e7-aa85-44dc-b555-33bb96a862db · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.061996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.061996Z digest=sha256:6bc87e7ad9149070c073fbeef1c3d0962592c3dc84739c585727ab1301706a7e

Observation a9fd87bd-9edf-4305-8e57-5b10c52f7099 · outbound

This paper cites Gonzalez.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Gonzalez

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.136896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.136896Z digest=sha256:6dc25de55daca98798db70a5f08982b23cae3083b3229cf767cd6ab32f7ecb51

Observation 5b103f33-1860-4f6d-b835-a1d09a3214f4 · outbound

This paper cites Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:41:31.164901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:41:29.211992Z digest=sha256:bf0cfff5e28c58803bb42e38c438028e2bfb5489c526569cb6f2cdfcbd1285e6

Observation 5192dda7-3a17-4409-a2a3-95e4236bd339 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.309212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.309212Z digest=sha256:6b07be77357d54feca46c52cbe14ced099c256db9cffd4a8de59da74f99cc8ec

Observation 14d24b52-b044-43cb-b105-0d52f8507a96 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.358576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.358576Z digest=sha256:485ddb5c992b280d2c1ec50f152ea1a41a54dbfd8b3c482df1ce6609184dd3b7

Observation a511c874-76ae-4ea0-85d0-2e683ec10817 · outbound

This paper cites Concise reasoning via reinforcement learning.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Concise reasoning via reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.407186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.407186Z digest=sha256:7aa1f2ae251e19cc582ebb605d7a48998b3d68f8067587f308920183823194ff

Observation 268f1b66-93e5-483a-b060-83623e1f2a68 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.485654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.485654Z digest=sha256:b8d35e691b65caefe35d1f6e21611f2ad947838d9a9e96f6f4b87b613b43c7e0

Observation b78ba4d1-5f53-4b63-b769-b4ea1b998dac · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.558645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.558645Z digest=sha256:e920c346ebe1b09332945379f10e012fc57a8ac24b216dd304698697643e2f0f

Observation dd7c4d8d-a154-4858-b2e0-37ecac62f55f · outbound

This paper cites Solving quantitative reasoning problems with language models.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Solving quantitative reasoning problems with language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.651369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.651369Z digest=sha256:9d55546bbdeca129041aba06ed290f4fbc3b8b26de29a0537a7e8ab5c89c07e4

Observation a8570983-096a-48d2-96c9-cb11b5ddd572 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.732447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.732447Z digest=sha256:f13d8f06867b1ba4ddbd5a876b9e3a06484f1757d8ed704ac09301507e5efb5f

Observation d83cab6b-6cbc-47a6-89d8-dc83d447fa79 · outbound

This paper cites Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems, 2024.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.811939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.811939Z digest=sha256:ffa11e7a6958a8435c80b029ab5aa522477dac4da1bb3873d22397c96d7a57d6

Observation 2ed5ca11-5da9-4e14-ae6c-686c2603a8b2 · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility, 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:41:30.939427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:41:29.893961Z digest=sha256:f609877487d2abe75c0a68ebad904ad2c9182cac38a20c11a50cda1517f29fe4

Observation 0be12f3a-3998-48ae-b6c5-b0585ec4b4a6 · outbound

This paper cites Qwen2.5 Technical Report.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:29.958468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:29.958468Z digest=sha256:0d57292bb2d791426dd0f394648d42b1476421235788f95c2ab37a3f8783fe02

Observation 03ada641-80cc-400b-98b9-b340269b483c · outbound

This paper cites Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.007684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.007684Z digest=sha256:d4692d4d3e7be2b1ec3a97fa1eb985aaeee83244ed9fa5fab8278cd3ddb502e6

Observation 267993ab-5ae7-41f0-9343-3a739f73831e · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.046501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.046501Z digest=sha256:06d749d10a9a6d783077f6565ac67f94dc85fb3ee6bcd28e97d5621fd138658e

Observation d2a9d0b3-0b5e-4d55-9cc0-2409c6bff85d · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.127471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.127471Z digest=sha256:f102301b1e1a2ee2b7a442ae0dbca16faf4f9b6fe5d237bce68dc0647709d3c8

Observation 2f985c0f-0f15-43e3-9239-8d282d54aa74 · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.176429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.176429Z digest=sha256:99bd82a19f6d888f28e8f6c4a9a6589c8dba06bbc0eeafaab364c6f34157ca5f

Observation dfa24d0c-3be8-487b-8067-7af43a5e905e · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:41:30.761315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T13:41:30.255652Z digest=sha256:caa44523e367368823baaf094bd509a6d28746ed6354c9da07dc111d51c367b6

Observation 21283f7c-89e6-4307-8f3e-8c67dc08414f · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.317826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.317826Z digest=sha256:ba383cf0d3752ae5b2472a4803665a834fcd279ca33c448d3a41caaf15b07ae4

Observation 166f3aa0-d6bd-4d93-93f9-12be5dd032c0 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.382317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.382317Z digest=sha256:33b1ba07509b94b8d9ec231c56fe6b8824e7298e26b039577f2c8956f67d448c

Observation a85ba105-22b8-4869-ad5a-b8688d9142cc · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:30.443487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:41:30.443487Z digest=sha256:09dd133e1eab1fc142ebe0661344418f184acc1b642cd99c2354f71d0ab8abd7

Pith citing papers

Observation 788f827d-c11a-4f5e-911b-d40096de2272 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:29:56.823186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:f638f12fa4dcb2ae3afc8b1425a27700071a84fb77b37c7d86fd86420c5d00ff

Observation 1a6a0e55-3b5f-4c3f-8143-3877a86ae2c4 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 165

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.656180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.656180Z digest=sha256:229a4e538861b0d81a6c02c73e0ba2539fdc81c948be89f4e4179ba5ef9e6bf0

Observation 8750ce5d-5a39-4973-b88c-f5a6ab0bb8e8 · inbound

When Less is Enough: Efficient Inference via Collaborative Reasoning cites this paper.

When Less is Enough: Efficient Inference via Collaborative Reasoning Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:43.015133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T19:27:04.267404Z digest=sha256:03948be8697c7bf62563e5ecb419fde6c4c40dac74cfc16140dc5251ddae090d

Observation db530fcd-f20b-452a-a13e-c6bf8b302f72 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:09.245378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:6f06f2efebd592fc069a466a8ee665ddf9ed8ac35bc128c6eb0704cc2ef8e346

Observation e7e74023-cbfa-40f8-bf5f-ce48eef83c14 · inbound

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning cites this paper.

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:43:05.960803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T06:40:06.103206Z digest=sha256:6bfcef8dc411b265b731d0e7aaed0389ddbee51bd391d2b0de10483eed480ee0

Observation 95d5fa2c-63b2-4dac-83bf-2b0e9d95bf91 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.877011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:76dd5b7380a2662b8099a13cb25ee3ebc7c3dd4f14cf1396036f7aedaa0f130b

Observation e6967e08-c3c8-4e23-8297-c4dd9a0ee001 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.087657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.087657Z digest=sha256:20ee15c1b68789f6b819b59da5028f5a820679ef51a126df45bc74d75726561f

Observation d3634661-fae2-4be1-8448-6bce7fdca490 · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.876796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:d78cc5b27de379dcccc637fb42aeaa60f8d0bb0bafdce97a5082ab5eee21ada7