Pith. sign in

Paper Citation Record · LEDGER

Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2307.02477.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.02477 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:26:07.791076Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc2dccdf-d5c3-42fd-9e58-6dd7403b02ab · inbound

CodeMind: Evaluating Large Language Models for Code Reasoning cites this paper.

CodeMind: Evaluating Large Language Models for Code Reasoning Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:55:59.724114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T03:53:55.964755Z digest=sha256:8ade2e291a10643d259714f749c5a1e4c365b017d2be1e0d4cde0a027ce0390b

Observation 62556a23-cace-46bc-8bd7-288500d2566c · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:34:43.061781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:3eb903b36c7ba15b1d1795a58c3b534dc6b440e734553a96ca067ecbb44015e7

Observation 72d5a5a4-9e23-4837-bf1d-6761112bb72a · inbound

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations cites this paper.

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:07.791076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:07.791076Z digest=sha256:f844eee4136c642af5ec0c1ee9dda7bc6f3b793f5bfd82332ca52fd3af443351

Observation 7c070b59-3f4c-41e0-b02b-0124c0af69c5 · inbound

When Do Neural Networks Learn World Models? cites this paper.

When Do Neural Networks Learn World Models? Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T22:12:50.234654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:12:50.234654Z digest=sha256:db7a6e818339f82ca7ce6c061418f15037710f4ed66d3a100d71c85dc5a9ffa8

Observation bb29bb42-4472-487b-8b80-686ea9d74086 · inbound

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? cites this paper.

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:12.114973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:12.114973Z digest=sha256:50aba32d983c05c6d7aa3a649af146feb62d87b10ce1a0a8204c880bb7cebe14

Observation 348c9fc5-24dc-48f0-9148-e211ba1f9b8c · inbound

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? cites this paper.

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:25.165193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:25.165193Z digest=sha256:bdb2518783258b137ae3f009a8a251698338a1259e3a24bbbded528262a6a80e

Observation 0f37452d-96db-4374-a55d-dd9ebd2a7fa8 · inbound

Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models cites this paper.

Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:00:57.202752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:59:00.400429Z digest=sha256:eaeb6d9d1989e035e6e428d21ed6abf147e4c960de6d063fe7983e23ce4393d5

Observation 934dde8a-37b4-41c8-8e43-43025358ee8d · inbound

Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures cites this paper.

Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:45:23.463395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:43:14.903173Z digest=sha256:63b3467b03c40c9ad64b5278f6e31118f8a6f5aa8407cf492486505745ac93a9

Observation 6f1e6424-c511-461c-b745-9f61eaa03f69 · inbound

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts cites this paper.

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:50:57.407403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:11:19.295354Z digest=sha256:f7b153bc1a0e6d2e890a5fd14908f34119a99f00099a3b15479ee07fd36e4865

Observation 5fa8a140-b3c5-4625-b317-12fc078b4972 · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.024087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:78007d08c9332fdbbcb5170b95e413cb8e2cd78e5a72cb57fb376707dde0fe05

Observation 6ef97197-df16-45ad-a0e0-1cd392c09bb8 · inbound

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination cites this paper.

DICE: Entropy-Regularized Equilibrium Selection for Stable Multi-Agent LLM Coordination Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.994598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T19:58:32.016341Z digest=sha256:754da6e238ce77ea2d55ab6eedd7ec42fb41cf8b39fce52b29e5b010885d6a8d

Observation 1733969a-8cd5-4930-9094-42d7c728d6de · inbound

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models cites this paper.

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:18.978102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:40:40.776789Z digest=sha256:a50323d875caada68a90bb649d04c9d95b69d85083e64dd97cc9a6b6cdc40a24

Observation f43aec6a-2efa-4036-9afe-fa959a6a4af9 · inbound

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds cites this paper.

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.684825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T19:18:43.558804Z digest=sha256:01d55347f71fe9846b5a92f8071243d60ebc2dea90b49e5f67bb78316b62c6ef