Pith. sign in

Paper Citation Record · LEDGER

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

As of 19 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 23 inbound Pith citation observations for arXiv:2505.11423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11423 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:57:42.432385Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:48:29.166327Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f90a6148-be6a-4cfd-923e-6a5223f375da · outbound

This paper cites write newline.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.296906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.296906Z digest=sha256:7873199b3022996b57862cc31d9fb33a0faade62cd1225a2bd15591694c8edb3

Observation 02d6ee03-dbdc-416b-87c8-57bf23427f18 · outbound

This paper cites GPT-4 Technical Report.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.301715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.301715Z digest=sha256:5bfd2735177f8c2ae67415b7d8fb167ae339b5ade0a2e1654dfb80619c3a5eb3

Observation ce271b57-27ef-4bfd-ae73-ae5af10aeaae · outbound

This paper cites Claude 3.7 sonnet and claude code.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Claude 3.7 sonnet and claude code

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.966822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.305683Z digest=sha256:9e39694787f96ac22edc34e24b0226ba38d89dcc6caee942ffc9a533f90122c6

Observation d8068035-37b6-4253-ae8c-e7cf702d529a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.309844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.309844Z digest=sha256:cc1a90955f1d5e24cc55ce19994f795059dd2a35425df99ea4e19ef425598c0e

Observation fdd3f9ce-e150-428c-8487-bcc27b70562a · outbound

This paper cites Language models are few-shot learners.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.953640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.313643Z digest=sha256:c2c56a389d3b77173454c53e67631c38ca265b57165408406978da9843f08e6b

Observation 7b29e637-250c-4737-9862-6c12d6337d00 · outbound

This paper cites Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.317612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.317612Z digest=sha256:fa99e36fc236d6246b7b00ea10201847150461cda2a1c7826b4fd7f517622da1

Observation e5918775-7177-4e68-961b-b16de8d2bd2c · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Scaling Instruction-Finetuned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.321495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.321495Z digest=sha256:15d0f24b605c16ba6afcdd598fa1f50729c8aebed55d7c1fe6f3b6dd06671521

Observation 6c087069-ea70-4ed9-8122-d0d0ea0c2d1a · outbound

This paper cites Complexity-Based Prompting for Multi-Step Reasoning.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Complexity-Based Prompting for Multi-Step Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.325240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.325240Z digest=sha256:7a8bf0d70383eb5cef595fdc3ecd21de0fde4fc49f72e4a11f6722c6da370073

Observation 37342762-b266-4b18-9002-dc49f4c9ca27 · outbound

This paper cites The Llama 3 Herd of Models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.329353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.329353Z digest=sha256:b8df44f99fb826d99869a5bcce21ad5675f03fc7971c493544801946bf0df358

Observation 491d17a5-6f4e-4d4c-9253-d231937784cc · outbound

This paper cites Algorithmic Unfairness through the Lens of EU Non-Discrimination Law: Or Why the Law is not a Decision Tree.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Algorithmic Unfairness through the Lens of EU Non-Discrimination Law: Or Why the Law is not a Decision Tree

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:57:42.748786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.332815Z digest=sha256:4b3350df7c1b35fff95206024e374cbe18b177a7f75444d27eabb75f7e5dc6fc

Observation ae3a7afc-466c-4897-81d3-025c667ec69d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.336482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.336482Z digest=sha256:7206e5662c41be1775330f195b0645eceb78be5cc81081c0967fb9b8ac08b83b

Observation a32da1d3-e84d-4e10-8293-141c0a7e9515 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.340517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.340517Z digest=sha256:ee91619d2595afa820d1a59b0d085e376dd9d1441b1c95ef071c39f1ddc2d6b6

Observation 1c20ee43-b435-4a45-9f41-a3fdf5340cc1 · outbound

This paper cites Do LLMs "know" internally when they follow instructions?.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Do LLMs "know" internally when they follow instructions?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.344933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.344933Z digest=sha256:d867b1c395be93b27b8f1ad163262734538e3253b076e4e4d3c615a08ab63cc9

Observation bc15cd24-b903-4832-bd9e-a88aa71d8bc6 · outbound

This paper cites Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-Text Rationales.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-Text Rationales

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.349215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.349215Z digest=sha256:3682b5feeedfdbc80fa2b1c5fb05247193f932f154e4c3456229f4da428e5171

Observation 4f635579-8978-4f89-a70a-dc28eec3cfaf · outbound

This paper cites Position: Llms can’t plan, but can help planning in llm-modulo frameworks.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Position: Llms can’t plan, but can help planning in llm-modulo frameworks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.353767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.353767Z digest=sha256:9bf1a61fe2f0a2e2d4f4c401221cfc9385deef8fd70bc1731bd9a769b100f722

Observation 0d88d199-8b7b-42ea-bf36-ca470545355d · outbound

This paper cites Metrology of weak quantum perturbations.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Metrology of weak quantum perturbations

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:57:42.667357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.357668Z digest=sha256:1338335466064bb606c9dff068f51aabed196836df38d397d8f6a1789cd6f47f

Observation d1b7c971-5168-4298-a2fc-bc96f6783a47 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.361500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.361500Z digest=sha256:7f4c04c484c85c169e56f1278d9cc058c4e2c26be0e6d04acfc2676bc76b3a6d

Observation 9f8d5873-075f-42a3-9818-892e8b75fcc2 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs The flan collection: Designing data and methods for effective instruction tuning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.934565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.365607Z digest=sha256:fa979d6d1169567c264bf766afdcdf4c5b1553e6391a6af3e110cfd7f1e65c21

Observation 967313ac-4c71-4cbb-90e5-c86893314412 · outbound

This paper cites Crosslingual generalization through multitask finetuning.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Crosslingual generalization through multitask finetuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.921626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.369850Z digest=sha256:66866543c160920ac2ec6407ff8779afaadfb9e3a24b306a9d34b538d03fcc52

Observation 2bda7492-7187-4d93-beea-2b31732383b9 · outbound

This paper cites Show your work: Scratchpads for intermediate computation with language models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Show your work: Scratchpads for intermediate computation with language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.374144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.374144Z digest=sha256:1c68ccd99cd842e21fecaa60cd41a166e4a526a99f6f55707f5a69d60c6b1850

Observation 4d478df3-9534-41ea-9f13-20b35d7d6949 · outbound

This paper cites Introducing openai o1.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Introducing openai o1

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.902245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.377977Z digest=sha256:b337013576a0f133c2a84be3d6c1cfbd6dc7bcca3a656e0c4acd7c17d724fba9

Observation a55ec9ce-6c62-4cad-b39b-42753f68db54 · outbound

This paper cites Training language models to follow instructions with human feedback.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Training language models to follow instructions with human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.381659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.381659Z digest=sha256:b602ce64a6ea90c3ff6df522429ebec5006ecc05cf0ee7e2c8a643603623f5b4

Observation 9d328994-16ad-4ea0-8fa7-a2eb1f762262 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.385847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.385847Z digest=sha256:14aca41fe797e844d161da6e981fddc1fa69f809aa20d8b69fa56fb8b964d811

Observation 74a1172e-28de-4dae-9393-70882d1ae31b · outbound

This paper cites Language models are unsupervised multitask learners, 2019.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Language models are unsupervised multitask learners, 2019

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.390038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.390038Z digest=sha256:93f45f8f09770fae94e849b7f4bcbdc01c3c119c982c6843934b59dee9d34b1d

Observation 65e23729-a29a-45e5-9fb6-cde33018e904 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.393794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.393794Z digest=sha256:9d5947b5dda797c4f5b677f49d622f21e730f1ed47e03d62d092fd93bbd846ac

Observation 74d28e30-eb18-4dab-a79b-2f59e3e88f20 · outbound

This paper cites Self-instruct: Aligning language models with self-generated instructions.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Self-instruct: Aligning language models with self-generated instructions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.876268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.397610Z digest=sha256:4fe8950170c60d4c040327cb36219970702bd86561c1c53ca57c3c858e11d3d5

Observation 96fc1174-f857-46a0-a0c5-23fda575c566 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.401269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.401269Z digest=sha256:f9f2e138c8e6e8858f703ea3d9474fcd3f7d0d33acac199c4ed9818d510787f6

Observation c1633b4b-9eed-4411-a180-59026c8659b4 · outbound

This paper cites Emergent abilities of large language models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Emergent abilities of large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.405252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.405252Z digest=sha256:bed623de5fe8a7cbe75d961d943df9a4d771fb217768911351577222bb59352e

Observation 85c62210-a548-4206-9f3e-2fdbf057811b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.408890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.408890Z digest=sha256:1cecb92ce733e7e6b36f70dc280e6967d7a4e50687a4a53f66acc9937ff78363

Observation 6c3b6a87-d713-4208-a50e-b99c48ac3477 · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Benchmarking complex instruction-following with multiple constraints composition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.837847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.412723Z digest=sha256:55659ea0670b3bac697c2475819026ac72312e718ca32537dc818107e2430067

Observation 91cbc755-63fc-49b7-8403-5e90d9fb4cfb · outbound

This paper cites A Comparative Study on Reasoning Patterns of OpenAI's o1 Model.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs A Comparative Study on Reasoning Patterns of OpenAI's o1 Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.416807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.416807Z digest=sha256:197209ecc2fe30ad87b9831b1d9f490486047a65a56b0ef3bc25fc78787df8c4

Observation 94f97e73-01b9-4af6-a7c5-9230ce42c9ac · outbound

This paper cites Re-reading improves reasoning in language models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Re-reading improves reasoning in language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:57:42.825392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:57:42.420831Z digest=sha256:a9525a4482e3eabdebc79cc59ffcb3d6b90293e35ba4e2ca6dd495aeda7ed763

Observation 215052a1-2c78-4211-a939-400aa24a684d · outbound

This paper cites Instruction tuning for large language models: A survey.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Instruction tuning for large language models: A survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.424526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.424526Z digest=sha256:8844f0d27a27642087c4561c250ae423f4de35e058abb2fc2bf7224c5958de49

Observation 0375dda8-faaf-4a1c-a1d0-04fff4f066f5 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.428434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.428434Z digest=sha256:f39bcca15075ed66c7ad58884d6a7ceefde3a9f24e523564d0684e7420415a10

Observation b00f4192-28e5-4c92-8a3c-1774c8988cff · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs Instruction-Following Evaluation for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:57:42.432385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:57:42.432385Z digest=sha256:96abaf7e6a31c407ea43ddf8495d6f96ba440bf4d23ce3eef00641af7ba03225

Pith citing papers

Observation ab12ca32-08b3-4294-b6b8-3b87497cfeaf · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.021927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.021927Z digest=sha256:7f3c29cdaf9d027c14a51cba5e5aeebe1a6cfa66d50fc5f76bb65bff490946c8

Observation eb443729-0a07-4c96-bcc1-fb66c56a6af4 · inbound

From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrieval cites this paper.

From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrieval When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:57:23.731384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:57:23.731384Z digest=sha256:4666d76c8164d5edfb8c4db5d2bb899026c0bb35daf55563e8713932b0153ea7

Observation 4d5a8e18-0965-4709-b7b5-188421683c1b · inbound

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity cites this paper.

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:10:31.638192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T16:10:31.440921Z digest=sha256:664f32ecb5903c63bbe530c7596ab20f028cc80ab4bc70d13c960eb542b87b36

Observation f6d90882-42a1-4a02-9b4d-b17aaff211e0 · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.568134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.568134Z digest=sha256:8b0710db6175d0ffcf128e5a9b7a4c5a2b1a585ce9f00a8b86d684e2cae4fe56

Observation 74af3947-339f-4281-8dfb-911079c03537 · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.171218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.171218Z digest=sha256:abb5fe2f42cce4ed079537186616d5f79131c1fc447b7c4ac9b03847ea318fdc

Observation 28d7d64a-746a-41c4-a2ae-0c8af1df3622 · inbound

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents cites this paper.

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:43:16.347192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:43:16.347192Z digest=sha256:e6941b3d4e5732c58214c515bbaea9fd02010542f925acb45523d49a5af8ee34

Observation 96659764-f58e-4e3c-8f14-3da85b10e709 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.758218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:1db35b97033789ddc9e9752cb622f0abedb6593b22138461713fbceb3d837043

Observation a18bf195-318a-4464-8997-06b27c88909a · inbound

"GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts cites this paper.

"GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:39.148228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:39.148228Z digest=sha256:d9e77910555772eed8f30610435c0b6d158da244dad468c9d66af78a5a19d263

Observation 8f11d1a3-2ba5-4015-b624-475e57d7c34b · inbound

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts cites this paper.

Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:42:38.706065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T13:42:07.883909Z digest=sha256:6ad3b1c9a1e2bd7ba239d766fc8d76f235a21a9cf088d188a90b7674e1c1821a

Observation 5912c812-75de-4777-bb03-f95516dfc9aa · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.400180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:ba90cbc5cf05db9373a8df2018e8af1e96186ff4e318f63142edd7e31e137a28

Observation c88f4463-35d5-4ac1-821e-ac6fbd29814c · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:48:29.166327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:48:29.166327Z digest=sha256:eedef473effa2109d1e9d65a3a9685ad6ffd5dce9938fb42f1091e276f626949

Observation f6d6b204-c636-4ec0-ac89-45171800693c · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.932278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.932278Z digest=sha256:60e6c191cf8fb6dde7e13966132d637f49924a6d726cd3993050e77710fdcffd

Observation ed3d7412-c7c4-418a-bb5b-199f77ac88cd · inbound

LightThinker++: From Reasoning Compression to Memory Management cites this paper.

LightThinker++: From Reasoning Compression to Memory Management When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:28:02.511545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T17:25:28.432170Z digest=sha256:0cd73935366222c1833169ae99370dce311442ee5e1f69c9dd914800ef84ed25

Observation 3c62f23e-7361-4d4b-a9b9-1d85788b7df5 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.525902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:a417d007abaa5003a97bde8ff0bdbd1a95f8679333635f4656e754dadbf78b89

Observation 15fdfa06-9bde-47b7-94b5-529da32323fe · inbound

DataDignity: Training Data Attribution for Large Language Models cites this paper.

DataDignity: Training Data Attribution for Large Language Models When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:10.356744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T11:53:19.594779Z digest=sha256:0004f63036b61158110b159273a5abc0ca795991cc476d3f075a5df65a26fb31

Observation e081aa12-7fc9-4bc5-b2ed-5278339ea8f1 · inbound

More Is Not Always Better: Cross-Component Interference in LLM Agent Scaffolding cites this paper.

More Is Not Always Better: Cross-Component Interference in LLM Agent Scaffolding When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.701864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T11:48:58.105919Z digest=sha256:65afe4bbd3744a31a03d365d16ad38d25130a40425a0e7ba976e5b7eb19a6d3d

Observation 4fc68512-6158-45dc-bcad-f86df5b9b6ac · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.363458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:2b704bdfa6be917c1cceab0e617f5559ae54251e810c821caae7711f168ee106

Observation e557b1c8-cf1e-4057-9d9a-c1152bca9336 · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:33.046329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:5b2c641067a193952562f28734a1af52b544128abff1254bf9ad623c8b8a8068

Observation 44e21807-6963-43e2-b65c-c109e9fd7e92 · inbound

When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following cites this paper.

When Built-in Thinking Helps and Hurts: Constraint-Level Error Shifts in Instruction Following When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:30.800737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:33:59.661761Z digest=sha256:437324c379103704c5e58216b1dc0ee66e7b98027649c9f5c9d4e7ab571efc38

Observation 207644e1-4782-47a9-a0b4-48865ba5fce0 · inbound

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models cites this paper.

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:18:03.816454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T09:36:05.700067Z digest=sha256:edde97fbbe4005af76c3439a070e12a174ead348c25abcec78c1f6d26efebc2e

Observation 603ce71c-60d2-4c17-a412-fcf6d5368fa3 · inbound

Structured Thoughts For Improved Reasoning And Context Pruning cites this paper.

Structured Thoughts For Improved Reasoning And Context Pruning When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T12:08:05.502310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T12:08:05.502310Z digest=sha256:f81c543061629808319ae03d167a901f14ccfc7814fb6d9d4e6bf1b39d3d1d86

Observation 7e9cc34d-0a3f-4738-a879-ee75b7144c61 · inbound

When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation cites this paper.

When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T11:24:37.140735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:24:37.140735Z digest=sha256:848cbc8d3271be0e6e755371c416b518c3c4f8bf143dfe3b446ea33bbac8cfd7

Observation e190c4f2-4725-41c1-a6ff-9c092fa92250 · inbound

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs cites this paper.

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:24.956984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:24.956984Z digest=sha256:b9b661c9e94c32414f981eca0caf1008bf55976ae3adbb1d6ca4a0e8e6984d8d