Pith. sign in

Paper Citation Record · LEDGER

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 8 inbound Pith citation observations for arXiv:2505.11827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11827 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:52:27.870796Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:59.391394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:13:58.683288Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved59
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e1dba22-f59a-42c0-bded-9c6f5d4bdb45 · outbound

This paper cites Learning to reason with llms.https://openai.com/index/learnin g-to-reason-with-llms/, 2024.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Learning to reason with llms.https://openai.com/index/learnin g-to-reason-with-llms/, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.848459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.634783Z digest=sha256:d270a7c014f20c53c93ee72e1860f2499fcd60081934ddd338ffbd35400a4e2c

Observation e3748f0a-bde4-4d82-9f1e-060c14ae6806 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.639546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.639546Z digest=sha256:50ab2b5f695a40aa7d878ae747661057db6128f06c8d67ac25fcdce5eda598c8

Observation 27ac5935-22e6-42db-bdd7-5244f26444ca · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.644098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.644098Z digest=sha256:16e1e1d663f925842ab23761de3f809544cd26803d5c71012775efca8b88529f

Observation a8f5373e-6d3d-4440-aa18-d314bbfad2f4 · outbound

This paper cites Introducing openai o1.https://openai.com/o1/, 2024.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Introducing openai o1.https://openai.com/o1/, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.647305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.647305Z digest=sha256:0e46687a247c7d2e88a55793f8431b2d895e590ee43bf2b168e683eb93f05ed6

Observation 6b5cfc72-4287-4bcf-b555-5bd139d88c34 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.650871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.650871Z digest=sha256:9c915bebc4ee130d1e543c8aecc6f443587ca312df3703dfb4c57bf331582709

Observation e2fc4f2a-5bb0-4c20-bbf8-728a31a1a393 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.658521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.658521Z digest=sha256:0e986198c9bea07d52c952a6735bc9f351c58e278327f2d1b4b073dc3a584173

Observation 2b157013-fddc-4f84-9a4a-d1d1a375fd2b · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Gpqa: A graduate-level google-proof q&a benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.661450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.661450Z digest=sha256:485ed9464c4f34a27e74d956a260af4cf083e4c05f80b1601e68ae27398c1583

Observation 2a969319-8f2c-4325-9f57-8bbdac892614 · outbound

This paper cites The Impact of Reasoning Step Length on Large Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning The Impact of Reasoning Step Length on Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.664401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.664401Z digest=sha256:5a6a0aa708b7b4d4700c96d867d73e3273e3931a08a9f78e72242fe938d9f5e7

Observation 31fb7393-e939-4dbd-883a-4a878431eb63 · outbound

This paper cites Compressed Chain of Thought: Efficient Reasoning Through Dense Representations.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Compressed Chain of Thought: Efficient Reasoning Through Dense Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.668099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.668099Z digest=sha256:d6a38f02a6b9c6139fa0a3e8ad9bf817837e0df4f8c6b490a45ea5cd0d8a9525

Observation 25566f2e-0ee5-481c-a220-65747c2038a9 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Chain of Draft: Thinking Faster by Writing Less

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.671402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.671402Z digest=sha256:a714c9774cd9e17577c5161bba3c7118d0addae5785f60a436f01fd3c8f7d883

Observation 2f4ef41c-fd30-4d41-9521-3e5d656545dd · outbound

This paper cites Overthink: Slowdown attacks on reasoning llms.arXiv e-prints, pages arXiv–2502, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Overthink: Slowdown attacks on reasoning llms.arXiv e-prints, pages arXiv–2502, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.822616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.675031Z digest=sha256:8b2a6789af58cc66a064d5db6f80935f486e7f5f905b201e33563eae097575f2

Observation f0b5292a-738b-42b2-8c1a-a7f835de7063 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Token-Budget-Aware LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.677819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.677819Z digest=sha256:ca317fc01377efa7c61f116683bfc6f5a29e7b48e674ef97ba3f9ab8bc2b254e

Observation 5b5bd5a4-9389-41e0-8ac2-56201a426bb3 · outbound

This paper cites Qwen3 technical report.https://github.com/QwenLM/Qwen3, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Qwen3 technical report.https://github.com/QwenLM/Qwen3, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.812993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.682141Z digest=sha256:25d3911c4780f79520a814bab386a91f873c1298287040cbf193e0e3dfe09386

Observation e908915e-3fa1-4ce7-bbad-d0357e9bf397 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.685481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.685481Z digest=sha256:b7e8d776b72198e53a079d04c36903aa1e631671a183c1d57741a3997f0cb445

Observation 2afa4310-dec1-4d01-a029-46e5ef2e3580 · outbound

This paper cites The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.688704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.688704Z digest=sha256:ec6c3f2b99c953e89fc61fdb6f656f36313f1a5b97233a8fd3c06ea7ea520ac8

Observation 850607b7-d2b6-4c07-9949-2d3547d9615a · outbound

This paper cites C3ot: Generating shorter chain-of-thought without compromising effectiveness.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning C3ot: Generating shorter chain-of-thought without compromising effectiveness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.692641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.692641Z digest=sha256:f6d53544d182e9deb072a343f7c27e9730bbf431e2060403d19620e9483e217e

Observation a84781a2-bf34-4873-9b11-d09c97742971 · outbound

This paper cites Dyve: Thinking Fast and Slow for Dynamic Process Verification.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Dyve: Thinking Fast and Slow for Dynamic Process Verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.695931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.695931Z digest=sha256:e53947950dda99fc9ed0e2f17db50c4cb3af083d8caddb6f09d4080ca1318ebb

Observation e6b09a1b-1fa5-48db-bfb0-559bfab9b492 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.699687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.699687Z digest=sha256:7e75f0d09607f0846abf5f9f95c65e1438018630216a55823837a910a037ceaf

Observation e6ff6136-31eb-4de6-ad29-1eb61a5528b7 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.704228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.704228Z digest=sha256:454442c0d724705061aa04e15a6642debdcc7b9ff6facde0d6182d8a5f54e3b2

Observation b20208bc-ead9-42da-a680-41fd3a7d5472 · outbound

This paper cites Online and offline reinforcement learning by planning with a learned model.Advances in Neural Information Processing Systems, 34:27580–27591, 2021.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Online and offline reinforcement learning by planning with a learned model.Advances in Neural Information Processing Systems, 34:27580–27591, 2021

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.798121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.707191Z digest=sha256:45d80c031f8cb8ce2acf32d8de59e98b73bcd5b3f5d6d76af920aba9542fe7b4

Observation 985d6bf2-2a28-42a4-89c4-d0e65b46cbf0 · outbound

This paper cites Dast: Difficulty-adaptive slow-thinking for large reasoning models.arXiv preprint arXiv:2503.04472, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Dast: Difficulty-adaptive slow-thinking for large reasoning models.arXiv preprint arXiv:2503.04472, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.710089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.710089Z digest=sha256:3a907825e44e3326fb103702f5856200cb4929f2024c9fee2d70de1f95b557df

Observation edff6d17-bcb0-4dca-a2d2-fb029ef536d1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.713053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.713053Z digest=sha256:8c2ca350fab2bcf98c5b3a0dd2e5dc1593fd2ebeb0a8aa508362b3b761ff86d0

Observation a1be2436-b5c5-4c4e-9c5c-4d84596db1d3 · outbound

This paper cites Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.715892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.715892Z digest=sha256:496845c2f3316d362dab0f5051017c2c58fccef48ab9e0d7544314c255400e0a

Observation 5986889f-1302-414a-ade3-7b66de653f7e · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.719429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.719429Z digest=sha256:72130827df57ab6af207934f53f89b008e66edbac2028949bd944c94aa62d9fa

Observation fadac48b-a9aa-41f3-8f7c-219aef7d1526 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.723462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.723462Z digest=sha256:033146bfe5285644b7d445b27b49f476c20a74df76c492005999605a80ad995a

Observation 4ba20099-7780-449f-806b-b73a8ac4f8cd · outbound

This paper cites ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.726983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.726983Z digest=sha256:14037262ba0a88ae1882fd4bc65ee9891162492933822c3a5bd390e51acb1bb4

Observation a40e80ad-2282-480d-8db0-e3cef605a3e0 · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.730903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.730903Z digest=sha256:8bd45c3c35a14f30046875e6db8feeb957cdb848051906b0e95a2698437bb027

Observation 765784a6-0dc0-4231-97e6-dbc6a8b6d504 · outbound

This paper cites Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching.arXiv preprint arXiv:2503.05179, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching.arXiv preprint arXiv:2503.05179, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.733962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.733962Z digest=sha256:47b9aa6a51ebd0585dff897f15397bd83302d65cb4acb5d4fdb4f3993e51142a

Observation e477dc24-4d10-4354-9d19-a314b218a1cd · outbound

This paper cites Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.736820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.736820Z digest=sha256:4837e7828ff81ed6a7d29fa11a92210d3d9ad206d6f6ae44c655df4296dbfe58

Observation 502911ea-43e2-4047-8336-6c8b367e5000 · outbound

This paper cites Multi-turn Reinforcement Learning from Preference Human Feedback.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.740448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.740448Z digest=sha256:ba7f10263bb234bcaf23585fe55983b96c9ab5d49578450bbe5feb9a4069518e

Observation 4ce1d644-c834-4bdc-b210-d44b430e6f02 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.744073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.744073Z digest=sha256:bdb48fb14ad27371b98aadd381637aea5d6cc1a4ec5e3124a7df3189f2f67013

Observation 31efee27-1b1e-43ca-9bb5-bd2bc54ce65f · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.747468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.747468Z digest=sha256:013d4cece103e91fdd42a216faca5e6d83848002dcaa4ca43ec14b5bed90e048

Observation ed3566b4-85de-4e3d-af50-03f48fbbac4f · outbound

This paper cites Understanding Aha Moments: from External Observations to Internal Mechanisms.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Understanding Aha Moments: from External Observations to Internal Mechanisms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.751482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.751482Z digest=sha256:06756e334f9501debcea1fa1e7a9d44fa6cf8d6051a4e44d286496f3f244d20f

Observation 47b0b348-7b01-4df4-930f-59b4a7fc3cb2 · outbound

This paper cites A mean value theorem in geometry of numbers.Annals of Mathematics, 46(2):340– 347, 1945.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning A mean value theorem in geometry of numbers.Annals of Mathematics, 46(2):340– 347, 1945

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.782955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.754781Z digest=sha256:eaf77b10219bb0449124d7dca4405c0d4b55e023ff891fd4a6ebbbe6ba9270c3

Observation fca55852-c6bc-4c30-9094-c5cea4f1f8cc · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.758311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.758311Z digest=sha256:e9e526204dd708a51704862f958d931f02f1d5f422ff761357b787565c583c78

Observation abb78ab0-561d-4ed0-9d23-7cb9bea99dd7 · outbound

This paper cites Group robust preference optimization in reward-free rlhf.Advances in Neural Information Processing Systems, 37:37100–37137, 2024.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Group robust preference optimization in reward-free rlhf.Advances in Neural Information Processing Systems, 37:37100–37137, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.762022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.762022Z digest=sha256:4f3b51a90e904b05ef74282b5a61d762c561ee8251b2bd205f7fa05f387cba75

Observation 27b61284-ff46-400b-9713-7e947da215eb · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.765050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.765050Z digest=sha256:2ec6a6de603b632bd69c68451c6d7a3a0df3b6153fef4a41c702de07ceb61b2b

Observation 520ef70c-bec0-4975-9246-5a424b32a98e · outbound

This paper cites The Llama 3 Herd of Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.767942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.767942Z digest=sha256:97fe3619f942eeaa591acbcd54441821be52ad374d727b8cf709bee4464814ad

Observation 618aa3ef-6adc-4710-ada4-e97572a50cee · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.771802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.771802Z digest=sha256:f356ff8af51ca1e718bb1f8d4c52a1ce79203f7ca4d995071af0018c7f0b87dc

Observation 7a23f9d3-94b7-438e-99dc-eb19b18c3fdb · outbound

This paper cites Aime problems and solutions.https://artofproblemsolving.com/wiki/ index.php, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Aime problems and solutions.https://artofproblemsolving.com/wiki/ index.php, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.766318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.775555Z digest=sha256:4ce740103e13fc01caf1b9e064795c6473190ad4a0e5af37bec52555a2c8bf5b

Observation dee9d6e8-c74d-48fe-9103-8c312fea37f6 · outbound

This paper cites American mathematics competitions.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning American mathematics competitions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:52:28.755841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:52:27.778876Z digest=sha256:45eac1d000691015a4b580a868c4b485083924f8fde73feff4b95a2383af81f4

Observation 655a4ab2-74dc-46d4-90db-9e1180f84fa4 · outbound

This paper cites Openmathinstruct-1: A 1.8 million math instruction tuning dataset.Advances in Neural Information Processing Systems, 37:34737–34774, 2024.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Openmathinstruct-1: A 1.8 million math instruction tuning dataset.Advances in Neural Information Processing Systems, 37:34737–34774, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.782764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.782764Z digest=sha256:01c89562a136a72c6912df19590c3dfe2424d71361c03534ae7daa1c1bae2246

Observation 4fe5c724-7d25-45a1-8fee-a2e28a7f8623 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.786159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.786159Z digest=sha256:7a1094870688b737f74e62b1254bc9cd86a413d583b4af161ac9a2210cf01913

Observation e802468b-9b07-4702-be93-1e2a8ccf773c · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.789583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.789583Z digest=sha256:28fbe9ae51a0608b18cd9750c517e0d04ccbb66329ea258dc82751df36136b43

Observation 94b89a96-f547-4572-aef0-f4973a882a20 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.792807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.792807Z digest=sha256:8c455e3be64325db592e4b6ec0d677f0f529a6d2e225d2f89b8c463dc76d5af9

Observation 14214f1d-8fa2-486a-afaf-38116cfeba96 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.796197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.796197Z digest=sha256:e8dc738d28fa95e55da0cbbad4319932558d1fd794f2bc80ac4d7cfeb6d02af3

Observation 574eb0f7-8cee-4c95-a05d-36c3acc8843f · outbound

This paper cites Reinforcement learning: A survey.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Reinforcement learning: A survey

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.799680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.799680Z digest=sha256:dc6440cf68a81f77f8ce89561e8e18e44093092f1553e42693610066c12f2f36

Observation be1a0a3a-9cd3-4480-9401-8af1e76688a9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.803424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.803424Z digest=sha256:930aef89975fe18233865394c7436497e2bf988bd50148a2e3bcf88c5320615c

Observation 46f35f72-7e74-4449-ab0f-915e23473538 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.806763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.806763Z digest=sha256:64a38537daeca30ca3a8c69594b8f9aa2ca30db8f281309f54b129f8a673e651

Observation 395f7013-2ea5-4df3-b813-2eeebbd629cf · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.809993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.809993Z digest=sha256:23b7e0a1545c5592799bacc8e4c6259dbd877ae70cbd24d2d0eb71c16d550ce8

Observation 184f5881-91fc-48fc-b1d4-f8fefd141967 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.812991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.812991Z digest=sha256:6abafc3c780be3a12a2e5629966e8683dec8cfe0341ed42d72e91045b895c18c

Observation 1742f3de-8276-4f9c-be08-59518fb6d091 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.816985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.816985Z digest=sha256:8559ed729c77240e3d48364db4c387b3ab431408d8cc43a7832f66a8fc8268ce

Observation 78f1c130-c44c-48b9-871e-a2aeea3c6cb7 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.820470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.820470Z digest=sha256:5e020694044fc872a51db8c94652aa2baef6bc14cdf351447e6be3284658ccea

Observation 776fd629-9eb6-44f0-b51e-a5c6e1992dcb · outbound

This paper cites How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.824350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.824350Z digest=sha256:ec2d668f1aac695ec38033b23b3a86445945b4a2ad96ab0d97794fa06362f466

Observation b7a02d69-6fa8-43ed-8bc6-e8f34317e1b0 · outbound

This paper cites Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.827471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.827471Z digest=sha256:112392827e30a7202ccbc1981b96cde02d2561d250163fc274a5b0d433205a18

Observation b3d3519c-2fec-46b2-87f7-24724493797f · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.831037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.831037Z digest=sha256:56c0da3368ad374bf51c9bfbf09bce04f8a90fb7e7d2707a2a5e9ce05f538c0c

Observation 7dd76593-d63d-450b-9f2d-31865105f96f · outbound

This paper cites Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.835260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.835260Z digest=sha256:d89b51c7feddde50c618f40d64b7bcceb53a1904615f22c61407eecc6aa75d08

Observation c9c5322d-43cd-4113-a0bf-1b7508694eb1 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.838341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.838341Z digest=sha256:bd6367795b8545d59392f24d58788f4afabc648649668e6ef52ef07383ac86bd

Observation d8b39363-d6ff-43eb-87ea-418ba1404c9a · outbound

This paper cites CoT-Valve: Length-Compressible Chain-of-Thought Tuning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.842063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.842063Z digest=sha256:fc6de0e1f9fabccb171ed6c8098b7a305f5a8e1442c7f9b0dfe52764b662162b

Observation 48c433c0-bc47-4fd9-a618-9f009a0397f4 · outbound

This paper cites CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.845785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.845785Z digest=sha256:1c5f4349edc9747c8c50471b6e0d2bff123c4eee5fff1eefdf33eec1a1bf8e14

Observation 109476f8-8def-475d-9ffb-b1d75da5e7c9 · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Self-Training Elicits Concise Reasoning in Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.849593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.849593Z digest=sha256:4cd96f5bb0d5cc0faee1e6f846a5f044f604878fa8a0b1c08de08598b051a5ba

Observation cca3f5df-7e8b-40bf-bd8f-0156febd332b · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.853828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.853828Z digest=sha256:7e19f26b4a865e51c688e507249f8691372edbcc74d795c733cd1450b9a65a77

Observation a821dc8e-b968-4b71-a921-b53b2085702b · outbound

This paper cites Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.856904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.856904Z digest=sha256:53b6e5ed24b35393102282ab18a425257071d12d0cb55f07631c09329a821f60

Observation 816da435-fdff-4de7-b8de-eb6a2dda2275 · outbound

This paper cites Lightthinker: Thinking step-by-step compression.arXiv preprint arXiv:2502.15589, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Lightthinker: Thinking step-by-step compression.arXiv preprint arXiv:2502.15589, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.860411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.860411Z digest=sha256:45591a2992ae2ecf49f1b312425c063aff8204bcbbbec0b4ecd45232f217ab22

Observation 12f1b8f7-0988-4533-ad2e-a02915e295c3 · outbound

This paper cites Inftythink: Breaking the length limits of long-context reasoning in large language models.arXiv preprint arXiv:2503.06692, 2025.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Inftythink: Breaking the length limits of long-context reasoning in large language models.arXiv preprint arXiv:2503.06692, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.864055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.864055Z digest=sha256:968fc95437b7baff36e53997b03095cafe2a4371b7b527fcc4f1fb34dda6a723

Observation a9e3d244-2771-4265-97ee-a3044ef22b02 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:27.866978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.866978Z digest=sha256:a502f06e7ebc2144e452c5fc08b106adca4f7c515b8089774ba367d9dd3d0a76

Observation 0f0da7bb-d5dc-4627-a65f-8716e3b0cd23 · outbound

This paper cites Given that47−1≡51 (mod 97) , find 28−1 (mod 97), as a residue modulo 97. (Give a number between 0 and 96, inclusive.).

Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning Given that47−1≡51 (mod 97) , find 28−1 (mod 97), as a residue modulo 97. (Give a number between 0 and 96, inclusive.)

Reference 68

Resolution
malformed identifier
no resolver link, observed 2026-08-15T20:52:27.870796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:52:27.870796Z digest=sha256:7c5f62f3f4a6d83159e05229095f927e03dc69ae3884d176f3958cf17146acb3

Pith citing papers

Observation c49ac434-a5a7-4105-9e34-bfe3de753a62 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.635307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:c56f52999f1de10fa5f73b4392d39c85b2a856dbd1eac89388c73e3043a75f71

Observation 54dd50c3-df24-42f5-b8f8-ec4744002cf9 · inbound

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs cites this paper.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.391394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.391394Z digest=sha256:b5c3faf3decd82fc81bbc42fd50250a8bc875fbd66de84b9506cf42381e7a667

Observation 8a4e6cdf-4396-4f8f-b656-829fb196fad7 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.525716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.525716Z digest=sha256:ddf98d00050e5c7e4f08fde90112893ef1028e3e2ecbf06e9bb4b3080e7a4a0f

Observation a8a8e0be-0c63-4936-a4fb-ca7f1444b804 · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.574670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.574670Z digest=sha256:7d1daf4816f221611f74781d349f20edcb80286458244e15760374fd87224b27

Observation 58067f36-3c1b-4b8f-9e29-0b375f8f95b4 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.128184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:df78b55a21ab9f9e8a76be8596c13a0b6a57da0bf996cf0f106d5ba2ba9ec4a9

Observation 2e4eb112-31b7-4a18-8ae7-d9998e1a7ce3 · inbound

On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective cites this paper.

On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.685140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T05:09:37.588841Z digest=sha256:962cecb1ca0191ebbc312b57bf6b9c648c1433b3c3a41bd44a83fe0736fd26a3

Observation 59fc4700-df1c-4469-b7c4-14d4160378d4 · inbound

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning cites this paper.

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:01.503753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:01.503753Z digest=sha256:069ef8d3dd3696159704a7c1b5bf7291dd5e8754931ed190942278a96de0b04b

Observation 0d5fba01-f539-49c5-bb77-0ff2e354f696 · inbound

Chained Recursive Language Models for Multi-Iteration Reasoning cites this paper.

Chained Recursive Language Models for Multi-Iteration Reasoning Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:18.195206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:18.195206Z digest=sha256:36cdc9dea342ef4be75c3a1bc907b02b3efe6e6ef6fb3de89e96d71773bd10cb