Pith. sign in

Paper Citation Record · LEDGER

How Far Are We from Optimal Reasoning Efficiency?

As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 2 inbound Pith citation observations for arXiv:2506.07104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07104 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:40.577890Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:53:48.051608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:22:55.270834Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved54
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcf89b3e-d763-4ed3-af92-b11c20922e51 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:34.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:34.499159Z digest=sha256:004662bc962907507d64f797e8fba3e96878ba9b65a8ba8efe75ba7e82aca5fd

Observation 8db244d9-d272-41db-af6f-7ea1ce392dfc · outbound

This paper cites Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025.

How Far Are We from Optimal Reasoning Efficiency? Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:34.596778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:34.596778Z digest=sha256:574ce1fd1a22dccbe615ec4bb61d91ddb4e6adbdcab4424a7222899d949c70f9

Observation 6938d49c-3201-49b0-89f0-480e540e7bec · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

How Far Are We from Optimal Reasoning Efficiency? HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:34.696226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:34.696226Z digest=sha256:6c9518eee0ce02f1c4fd1c70e7908e8d45fe31099bf62462f97b82b970c4dbdd

Observation 9769068b-d3e3-4743-a4f4-d51336ab5248 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

How Far Are We from Optimal Reasoning Efficiency? Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.028025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.028025Z digest=sha256:d338dfea4426c445d36f2ce607c252ff7fe48fcbb22aede25dad75bba252f0d6

Observation 24297eed-6dbe-462b-9d27-58e0c3675a5e · outbound

This paper cites Thinkless: LLM Learns When to Think.

How Far Are We from Optimal Reasoning Efficiency? Thinkless: LLM Learns When to Think

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.129268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.129268Z digest=sha256:b4fc7450cde94fc6479b9841e7962f2bdbef4b3d725d8d442f2d407b92fed9d2

Observation e92c41ce-5ba5-40f2-a2e8-addd6482a401 · outbound

This paper cites Efficiently Scaling LLM Reasoning with Certaindex.

How Far Are We from Optimal Reasoning Efficiency? Efficiently Scaling LLM Reasoning with Certaindex

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.259278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.259278Z digest=sha256:3194444272321b89f2af06db32eb2887705e0fba258c167dabe5003ef8e647cc

Observation e3fa284a-d61b-4d75-a819-64c533390610 · outbound

This paper cites Reasoning without self-doubt: More efficient chain-of-thought through certainty probing.

How Far Are We from Optimal Reasoning Efficiency? Reasoning without self-doubt: More efficient chain-of-thought through certainty probing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.358710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.358710Z digest=sha256:c8d07df69ba090ffe5ddf6065991376cfdb40e0482df52fee0d5abdf479cbc9e

Observation c1ffa9ca-0a8a-4e7c-9c76-86b37ecdcb0b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.468940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.468940Z digest=sha256:5e1bebb9cda01ea8ed02f5592f78653d01a9e6515575083047fb418c5d3a8404

Observation 640a7e38-dd98-47ee-bef0-a35c7b92c105 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

How Far Are We from Optimal Reasoning Efficiency? Token-Budget-Aware LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.589682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.589682Z digest=sha256:1297ac42fb9fe7080354c077e31d7f8d6406c6f0f9026a8ce3d841139f47bb78

Observation 9e3c7f12-3266-4a4c-a2fd-a26d3488b72c · outbound

This paper cites Skywork open reasoner series.

How Far Are We from Optimal Reasoning Efficiency? Skywork open reasoner series

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:42.383342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:49:35.713968Z digest=sha256:cfd7a5b523c4b6c92be8188d2cd6e732616c7757fddd9375bfe270a3effadd19

Observation 549e5200-69e8-4e89-9a22-a52ccaa7d237 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.815835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.815835Z digest=sha256:70d25cd171536a477af2ba072885f968ba06cde52cca02f1d36481335bc084c9

Observation acdd038c-6254-4ca2-82e8-ba764b5ad559 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

How Far Are We from Optimal Reasoning Efficiency? Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.897085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.897085Z digest=sha256:e7089e01aecdf00904fe9ac01777394dd67482d33a97081c287115cb4a2cb1d0

Observation 96c1265e-c539-45b8-b150-55341afd6eac · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

How Far Are We from Optimal Reasoning Efficiency? Efficient Test-Time Scaling via Self-Calibration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.050461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.050461Z digest=sha256:d05dd36820783dabf2d24de4b25035d7a2265abd29b71c32c1574a9eb40f1cdc

Observation 7b95b0bf-bfc5-4c36-a823-b52e471b94d0 · outbound

This paper cites Think Only When You Need with Large Hybrid-Reasoning Models.

How Far Are We from Optimal Reasoning Efficiency? Think Only When You Need with Large Hybrid-Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.120187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.120187Z digest=sha256:56968651db08dd0668cdcb019fb9dcb095b52e684de348ac3bd3a930d2d050b2

Observation 31a2eb7d-20eb-4a0d-a9e9-925851659732 · outbound

This paper cites C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness.

How Far Are We from Optimal Reasoning Efficiency? C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.187134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.187134Z digest=sha256:2f491ee3a136d928d28f9feca94feb3c4ebaf2e0f6a7b1f537071c0b8e5842ac

Observation d09a825c-5489-4883-9108-1d19d919c400 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

How Far Are We from Optimal Reasoning Efficiency? Solving Quantitative Reasoning Problems with Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.305178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.305178Z digest=sha256:3e80b188c854f57a029c2f8591fdff871aaee792d601ffbd7a29a3acc2bfb0c6

Observation 86b8cbe4-6afc-4a59-89fe-b898730394fb · outbound

This paper cites LIMR: Less is More for RL Scaling.

How Far Are We from Optimal Reasoning Efficiency? LIMR: Less is More for RL Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.467346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.467346Z digest=sha256:15b5574adae46cc6ec534e4a0d06e9a01eb8986246c1d5c28e22db91d897fe0b

Observation 2be6c0ee-0ca8-4e3e-abbc-00a4b245e895 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

How Far Are We from Optimal Reasoning Efficiency? ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.597126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.597126Z digest=sha256:877bf9759c58e3184c9571d75e88ea640ad29980ae30df6206156248045cc354

Observation 1a1662d7-9fd7-4836-a74c-7d216dbed935 · outbound

This paper cites Reward-Guided Speculative Decoding for Efficient LLM Reasoning.

How Far Are We from Optimal Reasoning Efficiency? Reward-Guided Speculative Decoding for Efficient LLM Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.717578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.717578Z digest=sha256:a444807fce15f555f534e83edc0f99ffb110f291b472ca1e1a7a12c6f6e07e9a

Observation 29cdb302-a971-4b8b-8d35-b197d55327d9 · outbound

This paper cites Can Language Models Learn to Skip Steps?.

How Far Are We from Optimal Reasoning Efficiency? Can Language Models Learn to Skip Steps?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.874631Z digest=sha256:6caa4a3896ba41032d85b7be66fda9cdb1cc3f8bc482b53be4d77c991198d8df

Observation 602394ee-2587-4421-b319-6b919e1f567d · outbound

This paper cites Fin-r1: A large language model for financial reasoning through reinforcement learning, 2025.

How Far Are We from Optimal Reasoning Efficiency? Fin-r1: A large language model for financial reasoning through reinforcement learning, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.987668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.987668Z digest=sha256:0a50f69856df9b212e168d5d0312fea028197a0e4d4878b108ca5adf84c6825a

Observation 58910aff-413b-4669-83d8-436907fc1c4c · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.109195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.109195Z digest=sha256:7d1f8470feaa8da781559b016167438d23ce430af0848dfa6c243bab16082df0

Observation 1ebbfb26-892d-4476-a44f-5c1083f4cc91 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

How Far Are We from Optimal Reasoning Efficiency? O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.215102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.215102Z digest=sha256:2eb9fcc71933221e3a520e27fc64dd0629fe3375facf91720857db70e68e2fef

Observation b0453166-e802-4e6b-93b9-798552d70296 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

How Far Are We from Optimal Reasoning Efficiency? Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:42.236363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:49:37.308823Z digest=sha256:83a1ae53bc80e89a73535a7402c9625e70947e22e253c5e48d3a250487b75c09

Observation c377da9b-340c-492d-9b00-f5ca907016bb · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

How Far Are We from Optimal Reasoning Efficiency? Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.401902Z digest=sha256:d6241a6b8cc3261209f7317390cc918991b72ef24c7914c8377d103742070a62

Observation 4597f838-16df-4b7d-875d-2b572aa48740 · outbound

This paper cites CoT-Valve: Length-Compressible Chain-of-Thought Tuning.

How Far Are We from Optimal Reasoning Efficiency? CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.540313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.540313Z digest=sha256:9d8375c19acc1fa11c5a0dc5927b2e95c49fd79f768da660cdea42adf104e567

Observation 7f0e3b78-4b8a-4a0e-b1d8-301a544ca41d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024.

How Far Are We from Optimal Reasoning Efficiency? Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.684760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.684760Z digest=sha256:b75bb9c00c223c3d26560e6b6e8824845a2085cf7741e103b566e7398409112e

Observation b0bfffb9-fc19-4282-bd12-c6c407233330 · outbound

This paper cites s1: Simple test-time scaling.

How Far Are We from Optimal Reasoning Efficiency? s1: Simple test-time scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.783145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.783145Z digest=sha256:5f083a4277118b66495ec3bf1882dbfe030383f91555f75c70a449e3130ec9cd

Observation 5fb48e33-8d90-4743-a1d6-bc5beb2fc1aa · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

How Far Are We from Optimal Reasoning Efficiency? Self-Training Elicits Concise Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.914205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.914205Z digest=sha256:2e006d4f109d0d8ed54ac418d4c684e159b8a33ba0dc80d5b96a475a164a88b3

Observation c885bb26-d3fa-4261-b1ed-9b76950c1eab · outbound

This paper cites Learning to reason with llms.

How Far Are We from Optimal Reasoning Efficiency? Learning to reason with llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:42.041410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:49:38.058152Z digest=sha256:bad32be24ce0eb2b7bddcd871f341730c7fcb663bc781c37dda8ac0e548c97b4

Observation 30c26e1a-4e14-4d49-b0b5-ba9f08e06c7d · outbound

This paper cites THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models.

How Far Are We from Optimal Reasoning Efficiency? THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.203391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.203391Z digest=sha256:fdeeb892e793678fa5f72e1822dfe57a65dd381d888906afeb5510c588e036bd

Observation 12d3f2f4-1a96-445b-97b0-651bbd5e1875 · outbound

This paper cites Optimizing anytime reasoning via budget relative policy optimization, 2025.

How Far Are We from Optimal Reasoning Efficiency? Optimizing anytime reasoning via budget relative policy optimization, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.275527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.275527Z digest=sha256:e04c8e7601354e79c888770422f20fae004bdcaee5215877f8e6e574ce3c194e

Observation 3b642e89-0914-4e15-873d-ee3d7fab022d · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

How Far Are We from Optimal Reasoning Efficiency? Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.378369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.378369Z digest=sha256:9967977238a34045d32a8bb63701a420733de0905e39186daee58b8ecb112ec3

Observation b79cc64b-1687-40a8-aa04-422e5c900707 · outbound

This paper cites Areal: Ant reasoning rl.

How Far Are We from Optimal Reasoning Efficiency? Areal: Ant reasoning rl

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:41.896795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:49:38.442171Z digest=sha256:c59f5f7260f93706ec39ddd3054f357cb0f643670af1f4e046eed9ce26c08e53

Observation 1a5c1ea8-2935-407c-8dd8-4cc7c720bb0e · outbound

This paper cites Hawkeye:Efficient Reasoning with Model Collaboration.

How Far Are We from Optimal Reasoning Efficiency? Hawkeye:Efficient Reasoning with Model Collaboration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.520326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.520326Z digest=sha256:7d39b11ce1292ab30290845288e3b796d8fa897dd5c7602bffea42e268c4ec6f

Observation 881669e3-6609-42a6-a6ce-41f8cef58788 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

How Far Are We from Optimal Reasoning Efficiency? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.588973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.588973Z digest=sha256:151daa9263a41263ba6ecfc8b6b5519140db3d0168d13572e3a03a3682740391

Observation da7c21a4-0088-49f6-9cba-0772561ccad2 · outbound

This paper cites Dast: Difficulty-adaptive slow-thinking for large reasoning models.

How Far Are We from Optimal Reasoning Efficiency? Dast: Difficulty-adaptive slow-thinking for large reasoning models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.672014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.672014Z digest=sha256:5a25bf61b9d47270fe7f55f8df5158284b0cfca8c1607e43310e9b10f7244f47

Observation f44f5bda-0b92-4924-abc8-e3dc9bcced9a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

How Far Are We from Optimal Reasoning Efficiency? HybridFlow: A Flexible and Efficient RLHF Framework

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.781973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.781973Z digest=sha256:9eea06e06907ab8fd625850a719aa43becdde872253856b4b05e8f4b91ebfd62

Observation 1ea4eec6-f2da-4d76-8c2c-e7ede3698bc7 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

How Far Are We from Optimal Reasoning Efficiency? Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.874482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.874482Z digest=sha256:43f825e7aa13b5b31621bc5308705cd5bb9868b403ccb52e65fe306d1cc0e64c

Observation 33a60b7a-8803-4660-9fd1-0d1350d14818 · outbound

This paper cites Fast Best-of-N Decoding via Speculative Rejection.

How Far Are We from Optimal Reasoning Efficiency? Fast Best-of-N Decoding via Speculative Rejection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.967912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.967912Z digest=sha256:83f978e75f2cb88e39ab69aefe1d48b986f116b246278e821f92dbe44d3313bd

Observation cc5e8e5c-9620-4805-8292-b4227d436795 · outbound

This paper cites Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233, 2025.

How Far Are We from Optimal Reasoning Efficiency? Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.034718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.034718Z digest=sha256:d81a60c187692d3b6b2504fa4b7e308bde2c6bf57de7b6a2086386660adb99ab

Observation 62567a75-a81b-4646-a49a-eddfb5bd2f03 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

How Far Are We from Optimal Reasoning Efficiency? Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.110881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.110881Z digest=sha256:8692a064b61bebcc050d6e335da81db2c3b36b7e519405a2076bad289a732be9

Observation c2579418-fe62-44ed-a47a-5b335be1e611 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl,.

How Far Are We from Optimal Reasoning Efficiency? Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:41.736320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:49:39.175313Z digest=sha256:d70b8d0e25673bd1d8bc658b87241bc6f9cde39fcc36ae96a37fc0ae54c772ce

Observation 71bb0eea-045e-40b4-be43-5834fba86634 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

How Far Are We from Optimal Reasoning Efficiency? Self-consistency improves chain of thought reasoning in language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.336776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.336776Z digest=sha256:3aa5585848018d00f89c51e336ec95136b8ccd03b14902a66a043f3497db9891

Observation f6eb5b31-48c2-4dc8-a33c-42347b1f8cf9 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

How Far Are We from Optimal Reasoning Efficiency? Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.427926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.427926Z digest=sha256:6bac1eb41208b9ca96f1cd4b9186375893563f37d2b071ec20a116ea3eec8ac4

Observation 50582151-68b1-42f7-aeed-02770b8720a5 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

How Far Are We from Optimal Reasoning Efficiency? Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.495400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.495400Z digest=sha256:147f62cbab3e089f12b8144daa911c655c24ad5c331358c3744902c2c4c0fc20

Observation f1a8cdf3-9093-43e4-b5df-b24053add6fd · outbound

This paper cites Scalable Chain of Thoughts via Elastic Reasoning.

How Far Are We from Optimal Reasoning Efficiency? Scalable Chain of Thoughts via Elastic Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.585269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.585269Z digest=sha256:e324b6f738ddb7d08cef66e32781835f1144bfded75bb99430c5481f0011112c

Observation 66b4e1d8-cb54-4e43-bcdc-9bba618053fe · outbound

This paper cites Qwen3 Technical Report.

How Far Are We from Optimal Reasoning Efficiency? Qwen3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.672165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.672165Z digest=sha256:6b6c011bbd8d982aff6f7e2b792aba72c280130a19638c78000ffe2d6d9ac080

Observation 69f3e7db-693f-42a6-82b8-c1ec8f1dc793 · outbound

This paper cites Think When You Need: Self-Adaptive Chain-of-Thought Learning.

How Far Are We from Optimal Reasoning Efficiency? Think When You Need: Self-Adaptive Chain-of-Thought Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.753920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.753920Z digest=sha256:a222c36117caa5cee8b2135952877e8a82316709c1ba6936a6e6f55081d69828

Observation bd9e2959-fb9d-4430-8ac5-4ba499f62c51 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning, 2025.

How Far Are We from Optimal Reasoning Efficiency? Towards thinking-optimal scaling of test-time compute for llm reasoning, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.840510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.840510Z digest=sha256:13a7f643134c201afccab124545bc41222afc2800b33c863fbd9f27f2a4a3efe

Observation 6af1f9e4-a779-45f5-a6d3-8a07b84ae273 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

How Far Are We from Optimal Reasoning Efficiency? Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.909550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.909550Z digest=sha256:54ac685428297c32471ade8bfbd7f341ca8ec89cba9fea6428c1f4a972dfeb69

Observation 589595c2-21f7-4948-a720-d00509f74398 · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

How Far Are We from Optimal Reasoning Efficiency? FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.978220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.978220Z digest=sha256:012a6e7ceb0ea546115045daa72665a46faa39d0790d664af41ce95b3d920f7f

Observation 3def0778-7059-489c-b93e-13253303564e · outbound

This paper cites Distilling System 2 into System 1.

How Far Are We from Optimal Reasoning Efficiency? Distilling System 2 into System 1

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.051297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.051297Z digest=sha256:7a10ab5e549863e3420b638063d8faf1399b255c0877f98ae8d567c69b405db1

Observation 6322f138-1f81-47ad-8802-2bdb1a21490b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

How Far Are We from Optimal Reasoning Efficiency? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.129855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.129855Z digest=sha256:6421ce20bac8b0ea2b4ac69f285496b415a4d41832619af6102551cf601215ef

Observation da1c3027-fe12-44b0-9033-217c568dacd6 · outbound

This paper cites Z1: Efficient Test-time Scaling with Code.

How Far Are We from Optimal Reasoning Efficiency? Z1: Efficient Test-time Scaling with Code

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.227034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.227034Z digest=sha256:698d2bc356c7b231bd59584c91f6d2d03a86b0e3a1406e5ff1689282deeade31

Observation 487a986d-c299-4028-900a-37dcdba052b3 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

How Far Are We from Optimal Reasoning Efficiency? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.307416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.307416Z digest=sha256:48f29215e7880700feb27ed7c3ea42a6150abc6f334cb752047a79a4c16c8b4b

Observation db605667-f43e-46e3-a2bf-6367e64205fb · outbound

This paper cites AdaptThink: Reasoning Models Can Learn When to Think.

How Far Are We from Optimal Reasoning Efficiency? AdaptThink: Reasoning Models Can Learn When to Think

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.388824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.388824Z digest=sha256:f928529c940eb9fa294889e6c317ac10e5d01ac9e5dbd491073d6f87f1ec52e2

Observation 2ebf69cc-e2b5-4f41-a2a4-bcc8a30e0fa4 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

How Far Are We from Optimal Reasoning Efficiency? R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.486336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.486336Z digest=sha256:2c5bac3053ac3abfd8155ef8e38d53562f4dbbf600caae450774bfaf308fa0c2

Observation b26ff1b8-67af-4eb3-8ce0-a5385bfacbe9 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

How Far Are We from Optimal Reasoning Efficiency? SGLang: Efficient Execution of Structured Language Model Programs

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:49:40.577890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.577890Z digest=sha256:7b869e194837c62f2b9ea438b54e2800d5008a1795601e88a4b778186e57f254

Observation 29e8489d-c029-4956-befa-11f8158d156d · outbound

This paper cites an unresolved cited work.

How Far Are We from Optimal Reasoning Efficiency? Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.249076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.249076Z digest=sha256:a48f28db24d233e5d378840426c85c127d72b3b2c234353b6441b92234e0c3d1

Pith citing papers

Observation c1c4647e-b274-4e3b-8244-f619c76f0c3b · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey How Far Are We from Optimal Reasoning Efficiency?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:48.051608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:53:48.051608Z digest=sha256:146dc83f8e80a43c8a8b56bb483c450d60493de99a95bc6e6edc5a7d031fcee4

Observation 8549c248-218b-49cd-b042-681bf3ad756c · inbound

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models cites this paper.

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models How Far Are We from Optimal Reasoning Efficiency?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:22:55.272653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T13:21:36.606855Z digest=sha256:9d7d4aa890323f4194002a61d63ba42162961227c237ef9d9ea6cf0356919941