Pith. sign in

Paper Citation Record · LEDGER

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 10 inbound Pith citation observations for arXiv:2505.22113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22113 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:17.785500Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:24.782504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:27:24.497080Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved21
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c384100b-c8b7-4d1c-b309-d55a9f9c7178 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.917705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:14.730171Z digest=sha256:c470fb8f045b9d7b57b7d49dfc950c986e7d1c47fd436cb4a081c3e126f75851

Observation ac1eb0a1-e109-406b-988f-f90903cd7299 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.651898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:14.874746Z digest=sha256:8220d107690727b5cff00f1d79a42dd545a28a80ae02fc97e6d58983502b5ce4

Observation 93ab8204-0496-486f-bde8-8f826a067208 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.425334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:14.990072Z digest=sha256:9d3fda4b7991e34f7ab560f440d5a3ed44b67729b510f0697cfeab508d21d6bb

Observation b24a198a-bec0-47a7-b87b-9f14c1907439 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.182153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:15.095136Z digest=sha256:0e9fd95f8abca26de4c23441b9976407ab3b5cf3078bff37b89b3da88017e544

Observation e93cbc2c-dc6f-4737-836a-fe532de576a4 · outbound

This paper cites solution1.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models solution1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.896976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:15.203468Z digest=sha256:fcc643371034eba01df05483480cafed260c14e05c47c8fd703fcccafcfb0c3a

Observation b0690b8e-de93-40a9-ab65-9a913a55ab7b · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.332107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.332107Z digest=sha256:8ff8167999675b25875ca1cdde5cacffa6101eb6f1e4a002d221cdac815ddfb9

Observation b94e670f-245b-4ca8-b275-77d94a51c5de · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 10

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:20:15.410585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.410585Z digest=sha256:06caebdcaff7f299cda2cdb3fed7d0e23f3427877c62c26ad12fae215e6219e6

Observation a0f55e05-4ba8-4f10-b81a-5675b489ce45 · outbound

This paper cites step_index.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models step_index

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.723877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:15.541578Z digest=sha256:a585f801d8d6ac0fafc5d4986fd4187527aeb39ef190b377188725e0d204f4e5

Observation 8d26cb6c-619a-45ca-9630-2699465649da · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:20.470923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:15.647111Z digest=sha256:f0c1de677b3892c533dc41758fc948750b02ad6a1f81bf25cf35ff3f6994d53a

Observation 03e0ef78-1002-4995-902a-e9e024f821a1 · outbound

This paper cites correct_answer.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models correct_answer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.257129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:15.714787Z digest=sha256:f273910ee32350d6e6702b4c46a607d6460593ba62c5762ece8dc52fc68ad8cc

Observation aef9f10d-7f23-4bb8-8aaf-f09ab6276f2d · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.844503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.844503Z digest=sha256:1ba1de6fe06e65dc8bd5146ea13cb0dce0a684a74b3fffb32587edd8fed0c961

Observation b14553b6-303f-45e2-8cb1-4b2373b2aad6 · outbound

This paper cites Match": Aligns with ground truth -.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Match": Aligns with ground truth -

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.984932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:15.969033Z digest=sha256:912ecc1a5ee17a866047e314501d1bc3ea18290070c257597b9c3c1645f5c2e6

Observation 8cbf3a86-79be-4057-ae6e-b8ff97544a75 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.007807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.007807Z digest=sha256:4da837e8f2cf5c853070645984b11f1c69f89fb1c29438d1f2bade4684767472

Observation fb0856b0-a0ac-4267-b70a-af6bcb28aae7 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.097596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.097596Z digest=sha256:bcb6131fdf7eefd3c2fc7c689ff1e68ae0d10b6a602cf0e3b8c42c9fbd408807

Observation 40bbb535-5b02-4e5c-af68-02dfae855d04 · outbound

This paper cites Always include the final step that contains the answer.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Always include the final step that contains the answer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.720710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:16.225473Z digest=sha256:fef32ca6cafc48233eba9b16fb1c25aa371ad49e7643ab04cb6e1f5b262a5637

Observation f4f114d5-2c3a-40ea-8af4-a05d4cd84ba3 · outbound

This paper cites step_type.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models step_type

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.432023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:16.407694Z digest=sha256:4c38b1e305e6a608180485c8a4099f40b82c4e801ec53d43e9a7b826c3f5a681

Observation 74c47502-d618-4397-8d1f-e3632be3252b · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505667Z digest=sha256:0eb999a363cc225156d04a160df539dbf33da73917d7651e6213e69a46775c76

Observation 839a257e-a783-4f7a-a736-932603356af4 · outbound

This paper cites 19 Invalid reflections include:.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models 19 Invalid reflections include:

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.156441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:16.674204Z digest=sha256:d8d8a6eb5c97ce489fec429da8648bc350bca667732225b5d8981d9515503aaa

Observation ad604f45-77b5-4b0c-819c-8c47d608303e · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.761555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.761555Z digest=sha256:02777a212446ebe9405b17105400ab1e359fba00c5bee35c07ed57d06d8d915a

Observation 838ee48d-2e6b-4dcd-8579-6b0d44f19a8a · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.841406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.841406Z digest=sha256:193e30e9b4dc5beedd7271afb54c56515f854d768e4e3651798d7ca2dc812b19

Observation 72863295-368d-4bbc-90bd-91da3ac42c22 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.934380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.934380Z digest=sha256:3541c608d4b6a98deaaef0461dc655712c8a7440df326d7b6b26ba21f46cd604

Observation 26672661-3b2d-440b-a9cd-7bf058668dfc · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:18.783558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:17.164581Z digest=sha256:df9a1a1525b768aa1164d43fb872742f434ae6fc438a9d96be1f908c95dcbc5d

Observation 89b1ac81-2722-4701-964c-bc09b596c553 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.242820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.242820Z digest=sha256:524f19d040a64832f0d4c8e9c04e2257bc69987aecbfb2da191fc137268fd439

Observation 92efe1ec-d154-4a5b-bb06-18470675cb3f · outbound

This paper cites conclusion.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models conclusion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:18.406192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:17.367067Z digest=sha256:2e84de58c4917907bff6971964cf7249cafaea2c682ac8f65a0ced20ab006e05

Observation db905805-387f-4bb5-8f29-99dfd5e8faa9 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 28

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:20:17.467252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.467252Z digest=sha256:064fb5bdbabb45577665a8a8e87040e5532f91a07efb2eebbcb0bd17f9c70d0e

Observation 12026156-5ed3-45a3-b32f-5f98b0d84f8e · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.535059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.535059Z digest=sha256:0cb2bfbffb0ee44ae531e3a64d684b200151f0a12c7f358182d250282a6131fe

Observation cc4b721e-f123-4f09-ad8f-ed4e19d84568 · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.616522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.616522Z digest=sha256:2dc2820eac61eff25a1f49c863fedd1895bddbb0a9336c0e718e204e1157bd95

Observation a50eea92-b63e-4e4f-b570-1d26d36d623f · outbound

This paper cites an unresolved cited work.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:18.118989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:17.785500Z digest=sha256:8c85e68576745a1e003709964d8a1a35b06bbad22ebbd0ff4776c55113d217ef

Observation 713c670a-8be1-4f03-a533-2e3109cfedac · outbound

This paper cites ARB: Advanced Reasoning Benchmark for Large Language Models.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models ARB: Advanced Reasoning Benchmark for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.608540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.608540Z digest=sha256:d9e4bab8f5e3a3fc43f68d48efd599b34afa125f8170a5da9940459a6f853eba

Observation b00ac0cc-38c1-4948-b687-b6e90501a4a1 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.394659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.394659Z digest=sha256:ab41c33d0c8969a46b2341c137e865f01bfb235cc89b165b8bce970b07d3fa46

Observation 9edd3401-5a7a-4679-a16e-ab99d28e7a5a · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:14.455229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:14.455229Z digest=sha256:b1efbd3dbd6b733819ccbddb61a52beaebb7893efc8563a2c16452b88895b0f8

Pith citing papers

Observation a4fcd218-24f4-4bd3-89f4-9ed46bc3b58e · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:24.782504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:24.782504Z digest=sha256:2e39cc3696aa429bbd7a5fe517524b58b10c184957e45a22cc9a938bc98f087f

Observation fcf09946-aa25-42e3-84b2-db2d72a9a230 · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:27:07.313600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:773caa1b5a7c975ed2b07341f8ba86e2e31d5641402d53500fffd4c60577c2f4

Observation 9f5055de-f5b5-42b8-8bc3-4c8f0321b50f · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.162463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.162463Z digest=sha256:984e93c43005ff08d95a2c3a2c4503072c9385bcf1eb03589e43f56311ec9be5

Observation 3753c99a-84be-4992-9af0-af1dace88f70 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:26.865491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:16014e7b2c6671af5e6748e98bf7a3f726c379c791c730a3bd67bfd3f867190f

Observation 823da83a-998a-40b9-9eb6-47c105ff82a4 · inbound

Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code cites this paper.

Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:06.021781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:18:42.321975Z digest=sha256:e2b435309a038c70095db1471de6e6d91ecdde2c3f7c929edb196d77ec6a3558

Observation f24ec194-6adf-4669-a666-cc30c21e49b5 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:09.086035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:01af2bd79e6f49f3abdbc50a446445361c1055cda05d3d5c2106f823adb3a1b7

Observation c7bc5be2-94be-4ac2-a0a3-5254c7bb8527 · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.557913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T04:45:39.641535Z digest=sha256:8936f305a899231cb3ed3a1022dd8e58b0faae952a0aa1e90c8a7876545018dc

Observation c28731c9-63ec-462c-a515-ebf21fbfbbb2 · inbound

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing cites this paper.

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:15:04.194995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:12:37.353935Z digest=sha256:b4add7c01626b35004939b7d57ce75a2a97662620d1ca5dada7a25c12afc692a

Observation 1243446f-35a8-45ae-9684-bb4a006b1d33 · inbound

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks cites this paper.

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.499140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:45:30.671490Z digest=sha256:695c8a3429f40c0d2043e97bce9c65e641a34dd6aa97057eff63a5bac8aaf634

Observation 864d426a-3310-4c57-a10c-2c3373f1cc2a · inbound

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs cites this paper.

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:24:32.223129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T09:23:26.386399Z digest=sha256:d574e5550bba3129786e0fe8aacb455a99b6b872ce714a47e7fe572ebce3c51c