Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:52:27.870796Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 8 inbound Pith citation observations for arXiv:2505.11827.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:52:27.870796Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:59.391394Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T05:13:58.683288Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8e1dba22-f59a-42c0-bded-9c6f5d4bdb45 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Learning to reason with llms.https://openai.com/index/learnin g-to-reason-with-llms/, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3748f0a-bde4-4d82-9f1e-060c14ae6806 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ac5935-22e6-42db-bdd7-5244f26444ca · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f5373e-6d3d-4440-aa18-d314bbfad2f4 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Introducing openai o1.https://openai.com/o1/, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b5cfc72-4287-4bcf-b555-5bd139d88c34 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fc4f2a-5bb0-4c20-bbf8-728a31a1a393 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b157013-fddc-4f84-9a4a-d1d1a375fd2b · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Gpqa: A graduate-level google-proof q&a benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a969319-8f2c-4325-9f57-8bbdac892614 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning The Impact of Reasoning Step Length on Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fb7393-e939-4dbd-883a-4a878431eb63 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25566f2e-0ee5-481c-a220-65747c2038a9 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Chain of Draft: Thinking Faster by Writing Less
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4ef41c-fd30-4d41-9521-3e5d656545dd · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Overthink: Slowdown attacks on reasoning llms.arXiv e-prints, pages arXiv–2502, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f0b5292a-738b-42b2-8c1a-a7f835de7063 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Token-Budget-Aware LLM Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b5bd5a4-9389-41e0-8ac2-56201a426bb3 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Qwen3 technical report.https://github.com/QwenLM/Qwen3, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e908915e-3fa1-4ce7-bbad-d0357e9bf397 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2afa4310-dec1-4d01-a029-46e5ef2e3580 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850607b7-d2b6-4c07-9949-2d3547d9615a · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning C3ot: Generating shorter chain-of-thought without compromising effectiveness
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84781a2-bf34-4873-9b11-d09c97742971 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Dyve: Thinking Fast and Slow for Dynamic Process Verification
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b09a1b-1fa5-48db-bfb0-559bfab9b492 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ff6136-31eb-4de6-ad29-1eb61a5528b7 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b20208bc-ead9-42da-a680-41fd3a7d5472 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Online and offline reinforcement learning by planning with a learned model.Advances in Neural Information Processing Systems, 34:27580–27591, 2021
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 985d6bf2-2a28-42a4-89c4-d0e65b46cbf0 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Dast: Difficulty-adaptive slow-thinking for large reasoning models.arXiv preprint arXiv:2503.04472, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edff6d17-bcb0-4dca-a2d2-fb029ef536d1 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1be2436-b5c5-4c4e-9c5c-4d84596db1d3 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5986889f-1302-414a-ade3-7b66de653f7e · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fadac48b-a9aa-41f3-8f7c-219aef7d1526 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ba20099-7780-449f-806b-b73a8ac4f8cd · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40e80ad-2282-480d-8db0-e3cef605a3e0 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765784a6-0dc0-4231-97e6-dbc6a8b6d504 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching.arXiv preprint arXiv:2503.05179, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e477dc24-4d10-4354-9d19-a314b218a1cd · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 502911ea-43e2-4047-8336-6c8b367e5000 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce1d644-c834-4bdc-b210-d44b430e6f02 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31efee27-1b1e-43ca-9bb5-bd2bc54ce65f · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3566b4-85de-4e3d-af50-03f48fbbac4f · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Understanding Aha Moments: from External Observations to Internal Mechanisms
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b0b348-7b01-4df4-930f-59b4a7fc3cb2 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning A mean value theorem in geometry of numbers.Annals of Mathematics, 46(2):340– 347, 1945
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fca55852-c6bc-4c30-9094-c5cea4f1f8cc · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Full Parameter Fine-tuning for Large Language Models with Limited Resources
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb78ab0-561d-4ed0-9d23-7cb9bea99dd7 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Group robust preference optimization in reward-free rlhf.Advances in Neural Information Processing Systems, 37:37100–37137, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b61284-ff46-400b-9713-7e947da215eb · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520ef70c-bec0-4975-9246-5a424b32a98e · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning The Llama 3 Herd of Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 618aa3ef-6adc-4710-ada4-e97572a50cee · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a23f9d3-94b7-438e-99dc-eb19b18c3fdb · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Aime problems and solutions.https://artofproblemsolving.com/wiki/ index.php, 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dee9d6e8-c74d-48fe-9103-8c312fea37f6 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning American mathematics competitions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 655a4ab2-74dc-46d4-90db-9e1180f84fa4 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Openmathinstruct-1: A 1.8 million math instruction tuning dataset.Advances in Neural Information Processing Systems, 37:34737–34774, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe5c724-7d25-45a1-8fee-a2e28a7f8623 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e802468b-9b07-4702-be93-1e2a8ccf773c · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b89a96-f547-4572-aef0-f4973a882a20 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14214f1d-8fa2-486a-afaf-38116cfeba96 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574eb0f7-8cee-4c95-a05d-36c3acc8843f · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Reinforcement learning: A survey
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be1a0a3a-9cd3-4480-9401-8af1e76688a9 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f35f72-7e74-4449-ab0f-915e23473538 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395f7013-2ea5-4df3-b813-2eeebbd629cf · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184f5881-91fc-48fc-b1d4-f8fefd141967 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1742f3de-8276-4f9c-be08-59518fb6d091 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ToolRL: Reward is All Tool Learning Needs
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f1c130-c44c-48b9-871e-a2aeea3c6cb7 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776fd629-9eb6-44f0-b51e-a5c6e1992dcb · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a02d69-6fa8-43ed-8bc6-e8f34317e1b0 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d3519c-2fec-46b2-87f7-24724493797f · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning When More is Less: Understanding Chain-of-Thought Length in LLMs
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd76593-d63d-450b-9f2d-31865105f96f · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c5322d-43cd-4113-a0bf-1b7508694eb1 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b39363-d6ff-43eb-87ea-418ba1404c9a · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning CoT-Valve: Length-Compressible Chain-of-Thought Tuning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c433c0-bc47-4fd9-a618-9f009a0397f4 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109476f8-8def-475d-9ffb-b1d75da5e7c9 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Self-Training Elicits Concise Reasoning in Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca3f5df-7e8b-40bf-bd8f-0156febd332b · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a821dc8e-b968-4b71-a921-b53b2085702b · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 816da435-fdff-4de7-b8de-eb6a2dda2275 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Lightthinker: Thinking step-by-step compression.arXiv preprint arXiv:2502.15589, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f1b8f7-0988-4533-ad2e-a02915e295c3 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Inftythink: Breaking the length limits of long-context reasoning in large language models.arXiv preprint arXiv:2503.06692, 2025
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e3d244-2771-4265-97ee-a3044ef22b02 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0da7bb-d5dc-4627-a65f-8716e3b0cd23 · outbound
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Given that47−1≡51 (mod 97) , find 28−1 (mod 97), as a residue modulo 97. (Give a number between 0 and 96, inclusive.)
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49ac434-a5a7-4105-9e34-bfe3de753a62 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 54dd50c3-df24-42f5-b8f8-ec4744002cf9 · inbound
MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4e6cdf-4396-4f8f-b656-829fb196fad7 · inbound
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a8e0be-0c63-4936-a4fb-ca7f1444b804 · inbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58067f36-3c1b-4b8f-9e29-0b375f8f95b4 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 216
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2e4eb112-31b7-4a18-8ae7-d9998e1a7ce3 · inbound
On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 59fc4700-df1c-4469-b7c4-14d4160378d4 · inbound
Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d5fba01-f539-49c5-bb77-0ff2e354f696 · inbound
Chained Recursive Language Models for Multi-Iteration Reasoning Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.