Pith. sign in

Paper Citation Record · LEDGER

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 10 inbound Pith citation observations for arXiv:2505.22113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22113 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:17.785500Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:24.782504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:27:24.497080Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved21
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c384100b-c8b7-4d1c-b309-d55a9f9c7178 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.917705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:14.730171Z digest=sha256:aede7338888793eb019c0647824596da9871794b264ff840fc91eb96261f05d5

Observation ac1eb0a1-e109-406b-988f-f90903cd7299 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.651898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:14.874746Z digest=sha256:3e7ac6f7120b242ea05aaee36aa66142637c3b875ed94b72f89eb3882f8f3cfe

Observation 93ab8204-0496-486f-bde8-8f826a067208 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.425334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:14.990072Z digest=sha256:725a1285ab2252686828a520709139d5d4e177d922394be784ad32a4e76ebb2a

Observation b24a198a-bec0-47a7-b87b-9f14c1907439 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.182153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:15.095136Z digest=sha256:7280873d3c21222e1ae3c6aa4005ff956c5518ad6846dddc60f250ec012d73e8

Observation e93cbc2c-dc6f-4737-836a-fe532de576a4 · outbound

This paper cites solution1.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models solution1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.896976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:15.203468Z digest=sha256:06603cccb68df2c05f443f4308b978d57543cee60ba7f30c9ac8300253f4d612

Observation b0690b8e-de93-40a9-ab65-9a913a55ab7b · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.332107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.332107Z digest=sha256:8ff8167999675b25875ca1cdde5cacffa6101eb6f1e4a002d221cdac815ddfb9

Observation b94e670f-245b-4ca8-b275-77d94a51c5de · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 10

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:20:15.410585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.410585Z digest=sha256:06caebdcaff7f299cda2cdb3fed7d0e23f3427877c62c26ad12fae215e6219e6

Observation a0f55e05-4ba8-4f10-b81a-5675b489ce45 · outbound

This paper cites step_index.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models step_index

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.723877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:15.541578Z digest=sha256:eafbde7ea386e7cbdf54ff0f85674d376dae4d191dfafd3095b0df77033ba1dd

Observation 8d26cb6c-619a-45ca-9630-2699465649da · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:20.470923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:15.647111Z digest=sha256:d52247cd23c438b23b57918bc5b583651f4ee8a04d82531da4d26081e7b8df22

Observation 03e0ef78-1002-4995-902a-e9e024f821a1 · outbound

This paper cites correct_answer.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models correct_answer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.257129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:15.714787Z digest=sha256:587951a7b30141c4e39600adef1362d063f267da4aeb45454a2ebe99c74bda07

Observation aef9f10d-7f23-4bb8-8aaf-f09ab6276f2d · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.844503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.844503Z digest=sha256:1ba1de6fe06e65dc8bd5146ea13cb0dce0a684a74b3fffb32587edd8fed0c961

Observation b14553b6-303f-45e2-8cb1-4b2373b2aad6 · outbound

This paper cites Match": Aligns with ground truth -.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Match": Aligns with ground truth -

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.984932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:15.969033Z digest=sha256:73a74181b1bca377f6a261cf7f85602eabf3e63dc40f8ab91fb8ddf41f7b4717

Observation 8cbf3a86-79be-4057-ae6e-b8ff97544a75 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.007807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.007807Z digest=sha256:4da837e8f2cf5c853070645984b11f1c69f89fb1c29438d1f2bade4684767472

Observation fb0856b0-a0ac-4267-b70a-af6bcb28aae7 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.097596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.097596Z digest=sha256:bcb6131fdf7eefd3c2fc7c689ff1e68ae0d10b6a602cf0e3b8c42c9fbd408807

Observation 40bbb535-5b02-4e5c-af68-02dfae855d04 · outbound

This paper cites Always include the final step that contains the answer.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Always include the final step that contains the answer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.720710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:16.225473Z digest=sha256:7a790e74c90c4b43c5dddbbc2e542217be9715af4747785fc92638e518f9ae7b

Observation f4f114d5-2c3a-40ea-8af4-a05d4cd84ba3 · outbound

This paper cites step_type.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models step_type

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.432023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:16.407694Z digest=sha256:042cafc5b41cd6a506e302ec2f04c3df38d0b4996b516318f4cd918231b06cf2

Observation 74c47502-d618-4397-8d1f-e3632be3252b · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505667Z digest=sha256:0eb999a363cc225156d04a160df539dbf33da73917d7651e6213e69a46775c76

Observation 839a257e-a783-4f7a-a736-932603356af4 · outbound

This paper cites 19 Invalid reflections include:.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models 19 Invalid reflections include:

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.156441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:16.674204Z digest=sha256:d7ca83ceaa4d990456d20b3878fed5295407a4c7a2b456f6895186a51a17ffc0

Observation ad604f45-77b5-4b0c-819c-8c47d608303e · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.761555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.761555Z digest=sha256:02777a212446ebe9405b17105400ab1e359fba00c5bee35c07ed57d06d8d915a

Observation 838ee48d-2e6b-4dcd-8579-6b0d44f19a8a · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.841406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.841406Z digest=sha256:193e30e9b4dc5beedd7271afb54c56515f854d768e4e3651798d7ca2dc812b19

Observation 72863295-368d-4bbc-90bd-91da3ac42c22 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.934380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.934380Z digest=sha256:3541c608d4b6a98deaaef0461dc655712c8a7440df326d7b6b26ba21f46cd604

Observation 26672661-3b2d-440b-a9cd-7bf058668dfc · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:18.783558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:17.164581Z digest=sha256:00b8efcb07ca6c2c90abfc3a9daf5488c681542ea6370db7165b16c4af751b30

Observation 89b1ac81-2722-4701-964c-bc09b596c553 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.242820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.242820Z digest=sha256:524f19d040a64832f0d4c8e9c04e2257bc69987aecbfb2da191fc137268fd439

Observation 92efe1ec-d154-4a5b-bb06-18470675cb3f · outbound

This paper cites conclusion.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models conclusion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:18.406192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:17.367067Z digest=sha256:aadc4c4d4e99940cf8a41eddb7ee0664d3022f9adb7567439cf836744eae5a2a

Observation db905805-387f-4bb5-8f29-99dfd5e8faa9 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 28

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:20:17.467252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.467252Z digest=sha256:064fb5bdbabb45577665a8a8e87040e5532f91a07efb2eebbcb0bd17f9c70d0e

Observation 12026156-5ed3-45a3-b32f-5f98b0d84f8e · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.535059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.535059Z digest=sha256:0cb2bfbffb0ee44ae531e3a64d684b200151f0a12c7f358182d250282a6131fe

Observation cc4b721e-f123-4f09-ad8f-ed4e19d84568 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.616522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.616522Z digest=sha256:2dc2820eac61eff25a1f49c863fedd1895bddbb0a9336c0e718e204e1157bd95

Observation a50eea92-b63e-4e4f-b570-1d26d36d623f · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:18.118989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:20:17.785500Z digest=sha256:dd83bc31eba0da8a7fb04657cef838bc42c9f7302fcd2bff0261ac41cf731975

Observation 713c670a-8be1-4f03-a533-2e3109cfedac · outbound

This paper cites ARB: Advanced Reasoning Benchmark for Large Language Models.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models ARB: Advanced Reasoning Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.608540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.608540Z digest=sha256:d9e4bab8f5e3a3fc43f68d48efd599b34afa125f8170a5da9940459a6f853eba

Observation b00ac0cc-38c1-4948-b687-b6e90501a4a1 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.394659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.394659Z digest=sha256:ab41c33d0c8969a46b2341c137e865f01bfb235cc89b165b8bce970b07d3fa46

Observation 9edd3401-5a7a-4679-a16e-ab99d28e7a5a · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.455229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.455229Z digest=sha256:b1efbd3dbd6b733819ccbddb61a52beaebb7893efc8563a2c16452b88895b0f8

Pith citing papers

Observation a4fcd218-24f4-4bd3-89f4-9ed46bc3b58e · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:24.782504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:24.782504Z digest=sha256:2e39cc3696aa429bbd7a5fe517524b58b10c184957e45a22cc9a938bc98f087f

Observation fcf09946-aa25-42e3-84b2-db2d72a9a230 · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.313600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:17aa336166f9c3c0609d9bb9b235052c2db474669b60c826a5fe24eb7b100bb2

Observation 9f5055de-f5b5-42b8-8bc3-4c8f0321b50f · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.162463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.162463Z digest=sha256:984e93c43005ff08d95a2c3a2c4503072c9385bcf1eb03589e43f56311ec9be5

Observation 3753c99a-84be-4992-9af0-af1dace88f70 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:26.865491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:db3d11552004b81f6289c190da63747fd043bfef7628ce4feb153cb2308b0ea1

Observation 823da83a-998a-40b9-9eb6-47c105ff82a4 · inbound

Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code cites this paper.

Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:06.021781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:18:42.321975Z digest=sha256:f883d5a2502d6b5fbe0f0f5161f6ea647b20eebb12b7484d4f3cfe3764d1f3e8

Observation f24ec194-6adf-4669-a666-cc30c21e49b5 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:09.086035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:7a73df36eeb9eb6e55637c3c31cffdd426f7d97c269a19f007c6fe3908be83dc

Observation c7bc5be2-94be-4ac2-a0a3-5254c7bb8527 · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.557913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T04:45:39.641535Z digest=sha256:e00429655e305aa2de32d6711f1553c736544fb6d8074861991811d6f8670e9d

Observation c28731c9-63ec-462c-a515-ebf21fbfbbb2 · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:15:04.194995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:12:37.353935Z digest=sha256:f13f59e00d0a4cabb323d861a1050d09594d067b73cd137271b428892f6d9dd0

Observation 1243446f-35a8-45ae-9684-bb4a006b1d33 · inbound

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks cites this paper.

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.499140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:45:30.671490Z digest=sha256:59f8cec1da791e8d15fe032c3eee6426e185a123f732e52a28b4413f89458c87

Observation 864d426a-3310-4c57-a10c-2c3373f1cc2a · inbound

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs cites this paper.

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:24:32.223129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T09:23:26.386399Z digest=sha256:e43303adb2f3d7c00bb38e61a44e5ca4dc82f31a6a3157a6deffb3176926db1f