Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2404.01869.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01869 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:29:05.198791Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:27:26.723998Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1afb69a-6236-47d2-8128-61707e539603 · inbound

TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis cites this paper.

TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:45:47.180451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:45:39.130509Z digest=sha256:8287011b6f51513d3ebef52ed3fb4999f03f99587db1b3f06ad184f43ddfbed6

Observation c436901c-6679-44ae-b83f-896fa2f89fb7 · inbound

Large Language Model assisted Hybrid Fuzzing cites this paper.

Large Language Model assisted Hybrid Fuzzing Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:15:28.694964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:13:13.613187Z digest=sha256:aa2241fb4d68aa7e3586d49945d83b770f993322287b65cd2dedfc165a63200e

Observation 495fd769-889d-484a-85a9-e2807d6b4a6c · inbound

Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment cites this paper.

Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T10:29:05.198791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:29:05.198791Z digest=sha256:f9e212f36c432f418daf22c1e3a12406ffed437af1bc2e63f3deadbbf5359217

Observation 18926fbe-dfa6-4174-bbb5-84dc97c31bec · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:47:15.149914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:ee084e96e12f9fd91dac68bc9f46ca530d3cb423e98c989331014e4738e7d9de

Observation 5eae57af-4b47-48db-ac84-34593590a22b · inbound

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation cites this paper.

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:51:08.241334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:51:08.241334Z digest=sha256:394418eaa7516ded6ad293078d2d7757a4c51265a8f0a4c4172fee44a13ec0c9

Observation ee5f1057-1385-4871-b8bd-00f9d536a1e2 · inbound

Propositional Logic for Probing Generalization in Neural Networks cites this paper.

Propositional Logic for Probing Generalization in Neural Networks Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:57.700850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:57.700850Z digest=sha256:c4d3ce00006f568ff0acb93004be0ff9282ebfc5e1bbce535eba8d81e0bc93bb

Observation 3608d722-55aa-45b5-8579-e4e6c896ec9e · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.602021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.602021Z digest=sha256:a506348bd2b46d7562c7732696b62db47699d8b1e1c3ebe30a80fc898bc030c6

Observation 0665d821-a98b-4193-949d-93d704d56573 · inbound

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? cites this paper.

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:23.526141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:23.526141Z digest=sha256:5b7d63b9722d4356a1e16a4c708ae77d7d832adc65a195cfaed556ba9c263e51

Observation c30afdeb-77fb-4f63-aa54-b3bce31efbbd · inbound

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models cites this paper.

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:22.769487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:22.769487Z digest=sha256:83449b105561b93a160e03edf2d21f032c68fcbe24d35023c689396f75cab242

Observation 0201270e-fd6e-4c99-a112-338a45d35d92 · inbound

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies cites this paper.

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:36.232363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:58:36.232363Z digest=sha256:966ff53f739c08e368165a143c9578f9306c2a0a475c6e291009f59a1c05a7fd

Observation 8d2d8d73-ee8e-4180-82a9-32fcbcc1dac6 · inbound

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models cites this paper.

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:42.582894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:42.582894Z digest=sha256:635282e5498d86c76963bd7472d22148a7ba10705bef0e393e98e3226a4af59c

Observation a9029285-c2f9-4d4c-82fa-3c3c26941291 · inbound

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation cites this paper.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.094226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.094226Z digest=sha256:7141745d43e8f429443e9357de04a0676d7e99b792fd23848a6b15c1f5728f93

Observation 3f16f27d-be57-49ad-8229-e64dba374242 · inbound

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue cites this paper.

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:46:53.218414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:46:53.218414Z digest=sha256:9cbc192ddb7115e96f1c82535cfb96b689e594f5cf38bed4587bcf230af1565a

Observation 18f575ae-5a69-4698-a01e-48370b28e9fa · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:55.547037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:55.547037Z digest=sha256:b730bc81038d17ca41488c17fedfbae8559472bdd740610b7d9a01834505cade

Observation d019ff42-c5a4-4de4-a718-a161f8264ee1 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.358111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:2674e0a0697f5f0ec1e40095e56ccace3d32568102d441639110aa2f3b53b1f5

Observation e57b035f-3313-4c7f-9912-7c9a63b2e998 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.419676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:f1a98d04a05aa0b6e64adc01896693dfdc78acd54a910a4eed187212e801b97d

Observation 5e95860f-12fd-4d31-8e4b-8cc8a0e2b51c · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.704850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:da920ebbd0fec824dc196d257a78471e4f8f28fa16232283cd00fe684e6d79ec

Observation 5477db6f-4f37-47e1-9e7d-fbf2e56e0112 · inbound

Reasoning emerges from constrained inference manifolds in large language models cites this paper.

Reasoning emerges from constrained inference manifolds in large language models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:01:15.379951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:58:30.713839Z digest=sha256:9b010aa65d18bd0066431570fbf1eb7ef480cac7abd121b47e7d6ea0fbb9d908

Observation 2700da80-4028-4f46-96d9-30ebecdca0d8 · inbound

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency cites this paper.

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.575582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T17:31:39.958584Z digest=sha256:de8403ef19125c6ae08bea7fca10477d3fd91b9e95380febf15c9882fe686d9e

Observation 2af6ec53-3384-4746-b4df-6e62da89311e · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.106736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:8cc0b4da2e66294a89b952b9379211b6f5e8900a202446e93dc77da4c4a20f2b

Observation a18e703d-5e5b-4251-957c-40f7849ac7ab · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.593743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:2f66c00cf31fee59e92a12ee91801be850a453247c55b4fb88a09ac4dabe9330

Observation 480fcf42-782a-47c0-903a-d240184b144c · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:05:31.924167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:1a08b4976fefb5811e822654cc6d465dd7129d66070d5c1ce3335aa31697a594

Observation 157347d4-b946-48ce-b33b-6149dfa131a4 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:26.725939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:2f5a4235dca59ff6e7c86d88db9e19cc5a246e277aead5b9949d202d4478b944