Pith. sign in

Paper Citation Record · LEDGER

THiNK: Can Large Language Models Think-aloud?

As of 10 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2505.20184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20184 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:37.575936Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:53.998837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved15
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a7c8880-5fd9-469d-bcb4-c7d5e50c4e4d · outbound

This paper cites By analyzing which concepts are invoked, we assess the knowledge dimension activated during problem-solving and whether the LLM navigates these domains coherently.

THiNK: Can Large Language Models Think-aloud? By analyzing which concepts are invoked, we assess the knowledge dimension activated during problem-solving and whether the LLM navigates these domains coherently

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.254905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.403974Z digest=sha256:b9917d5889896666414bfadd47d7180ecbc857ab27c63e10809123d9a31ce5dc

Observation 0fa765ff-d851-481b-a821-0373e1944790 · outbound

This paper cites Thilo Hagendorff, Sarah Fabi, and Michal Kosinski.

THiNK: Can Large Language Models Think-aloud? Thilo Hagendorff, Sarah Fabi, and Michal Kosinski

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:44.055481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:34.621929Z digest=sha256:4cf9949708014b7e8441abc509c2a578d7993542c5f99173026876059032323f

Observation a4aa73a5-1796-43c5-aa1f-a76b5d6ed06c · outbound

This paper cites Representations are critical for logical coherence and traceability in reasoning.

THiNK: Can Large Language Models Think-aloud? Representations are critical for logical coherence and traceability in reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.892838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.570182Z digest=sha256:4f395b233a35cf73f3a003b985e5c8e5b8c7103b191c32421cc9b077fd94969f

Observation 2f68695c-0270-47e6-81cb-67b468eebf67 · outbound

This paper cites A model’s ability to adapt its reasoning across such variants reflects generalization ability—an essential attribute of HOT.

THiNK: Can Large Language Models Think-aloud? A model’s ability to adapt its reasoning across such variants reflects generalization ability—an essential attribute of HOT

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.649080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.638056Z digest=sha256:cbe47e81762286f95a00b803bd93ccff2c35ad23acdcd082a0b0ca929bfa6178

Observation d9d1a54d-e351-4c68-a515-5b41cb3329af · outbound

This paper cites Five Keys.

THiNK: Can Large Language Models Think-aloud? Five Keys

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.396321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.760718Z digest=sha256:b3c1ec3376f281c37c73798a08c06023175e2dfc2295f4832319daca1f3a04fa

Observation 30a63fb8-e0f4-41a6-a468-e77e64a0ac3f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

THiNK: Can Large Language Models Think-aloud? Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.989895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.989895Z digest=sha256:4059faf9c3cd70ada7d044f5b7f1b2b105823a1e3cdd385429bcb315339c2665

Observation 6f285043-ceae-4a7e-b7ac-a08237cb27f7 · outbound

This paper cites Assessing and Understanding Creativity in Large Language Models.

THiNK: Can Large Language Models Think-aloud? Assessing and Understanding Creativity in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:35.188891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:35.188891Z digest=sha256:07c1af94a5e3737e6661ffb56fba29a867d4625c0ba026f100ff8e9bc525edaf

Observation eda686bc-321f-44f8-b9d2-085ef43b9912 · outbound

This paper cites Five Keys.

THiNK: Can Large Language Models Think-aloud? Five Keys

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.413353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.274686Z digest=sha256:e41eac88bb34a37ce9eb4c17eb287e0d84073ae3dd7b4068aefda1a4b935a0ad

Observation fc6c6ce0-d87c-436f-b349-e6e86497d3cf · outbound

This paper cites These skills serve as proxies for prior knowledge and inform whether the LLM draws upon relevant background competence.

THiNK: Can Large Language Models Think-aloud? These skills serve as proxies for prior knowledge and inform whether the LLM draws upon relevant background competence

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.501537Z digest=sha256:8aa4e3beb4d7e236235bebc0a6cdfe0964ef4a565746759586a367bbab1d36d3

Observation bc43f647-0c80-4f11-92ad-366e129e0133 · outbound

This paper cites You should try to understand and retrieve the specific mathematical information in it such as facts, patterns, objects, or contextual information, and decipher these meanings.

THiNK: Can Large Language Models Think-aloud? You should try to understand and retrieve the specific mathematical information in it such as facts, patterns, objects, or contextual information, and decipher these meanings

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.080713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.856691Z digest=sha256:3321166954b82779e456846d28f28893233519569fee28da8a29dd90e47b9936

Observation f80c2551-b0ae-441d-87f4-f27be7b9cd08 · outbound

This paper cites It includes understanding and organizing information, analyzing relationships, drawing conclusions, and distinguishing nuances.

THiNK: Can Large Language Models Think-aloud? It includes understanding and organizing information, analyzing relationships, drawing conclusions, and distinguishing nuances

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.709139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.926296Z digest=sha256:b1742b40acf39f85e5a86e38983adbead92e4d60a3f1dad00e086c1c4817ab6f

Observation 632569ac-4052-4187-86a6-dc67e329e041 · outbound

This paper cites These new expressions should have the same form as the given expressions in the previous generated math problem.

THiNK: Can Large Language Models Think-aloud? These new expressions should have the same form as the given expressions in the previous generated math problem

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.478200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:35.991143Z digest=sha256:cb34daf2a82294903d4f5ef81b9baa12de76fdd448be8a316e063effe285f08d

Observation d90fac4d-e73b-47fc-bc1f-49e61c612496 · outbound

This paper cites The generated stories must be a mathematical word problem with the corresponding expressions.

THiNK: Can Large Language Models Think-aloud? The generated stories must be a mathematical word problem with the corresponding expressions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.250540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.112437Z digest=sha256:388ed914383f06244ae0d8b7e144355a2ccd93eefd0e2fb5d8d8e3e3ff0c35e6

Observation 4ae966b1-69d3-4983-b64b-2753e21a22c2 · outbound

This paper cites question.

THiNK: Can Large Language Models Think-aloud? question

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:40.942421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.212697Z digest=sha256:afe7113286f4ed805d76fc246638bc22777a5b64c165bd3cd376a934cf73803f

Observation d7221b4a-69f1-4292-9c7e-b07d16682968 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.731850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.299113Z digest=sha256:9ff1800384260ec4677a64c490ed03baa5226d6905bfafaf914ac7d5b1ffd6d2

Observation 097e0f7a-e8c0-4aa7-8a53-1112802d3640 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.516278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.371403Z digest=sha256:dbf868df8109cdf5dbc641b83c8f2d0ee4d8c766ec2b40d7c07b68540ffc54d8

Observation e7cca71a-fa57-4018-81a1-996ea4f8a584 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.346332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.467020Z digest=sha256:e2a0e281b6166c81d7c285041aff78f2ce45f37d3042c21423741a7135cd4c20

Observation 71cfc3bd-c73a-4638-a83b-201870d4932c · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.119981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.557122Z digest=sha256:6cc744b3edd5674a508d1e95a8af7f3fe99933b0874a35fd5dcaa1e89bf81d0e

Observation 0bd16435-849c-4b77-b423-35024e0a2a0d · outbound

This paper cites ID": null,.

THiNK: Can Large Language Models Think-aloud? ID": null,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:39.859244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.650966Z digest=sha256:7c39cb419872052384ee189a91f8040782732239a3e090ccb2f4398c79c8d60b

Observation 905ea987-80e9-4efc-ac45-90783e9ccf18 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.683488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.725952Z digest=sha256:464934300d5c67a965cfac651d6376fcd441e7e13744a38d33bccf6752216cd5

Observation 373900a1-c4ed-414f-acab-352e4fb4840a · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.422480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.802182Z digest=sha256:33b98a4ae532c4d643f1d908a5df76d3909d6dd91775631493701307c6065fed

Observation 28bbc1f8-ccd1-4e0f-a597-8346676859cb · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.228336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.871190Z digest=sha256:95f23fb79a083c3f4085b8e2ece948a7e8952c374a85dc8057c423942eda94f2

Observation 79f727b5-a1b4-4900-b207-266a76810f5d · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.013489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:36.981345Z digest=sha256:c3e65c0bde7db02c013d23abb2c61dc027a174e254cb4de44abbaa5a3f8538d7

Observation fba5fb64-1ef4-447e-8f46-3851e00a2fc2 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:38.843275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:37.076085Z digest=sha256:d5d6d27068de530ecd50f02f13c9bd22dff2b6ba17be49b268321ae25c73fb3e

Observation daaf974f-20f5-4e54-9b92-d6d3073a526b · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:38.604568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:37.196578Z digest=sha256:1571487597765f52f49d929e2d0f342dd13302df00e491a4c25fccd29ad53a39

Observation 88110793-a66b-414c-aaa9-dd9f02d9fef1 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 31

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.445619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:37.286441Z digest=sha256:130b4bf6dd26e519a30110742231cd53298346c381f219191024319ccea804cc

Observation acfc7e07-2c28-4479-9c14-bb5374b74b76 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.191305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:37.367133Z digest=sha256:54a8f2c4eb1344f85d0c08423eec4b3322aa974444a0daff352f02dcc2e89171

Observation 237518b0-cf69-410d-8cfa-93d6a34f5c0c · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 33

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.000547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:37.458862Z digest=sha256:5f4302a12f577a6787414f8c30edf27470583b87c39eb2f19ebade5c255b41b7

Observation 7735d87f-a2ae-410b-84eb-b1424ad23de3 · outbound

This paper cites performance_score.

THiNK: Can Large Language Models Think-aloud? performance_score

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:37.798360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:37.575936Z digest=sha256:d8209c5dd87250ce81026605a8f3cf87406183db76f1096d95f8fddd1b8a0eba

Observation 6e3bc5a6-54ea-4d38-a129-9d44eba41060 · outbound

This paper cites English language teaching, 3(4):237– 248.

THiNK: Can Large Language Models Think-aloud? English language teaching, 3(4):237– 248

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.595243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:34.885086Z digest=sha256:a0def03fd78832a176949fc1ec85ef0ad270f353ed64c1a18a1d3b9ed497fa06

Observation 4f4a0605-e26a-479d-bef5-6bda3f3c3127 · outbound

This paper cites MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models.

THiNK: Can Large Language Models Think-aloud? MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:35.106825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:35.106825Z digest=sha256:5a3149c0bd3a18866920629ad4453c7dcea72126626f3edde8b691a85c164a7b

Observation 93bfd8cb-0986-4884-9228-76953960d739 · outbound

This paper cites HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.

THiNK: Can Large Language Models Think-aloud? HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.719074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.719074Z digest=sha256:41558339b954b416c8bacbd29f558c824b05d59eed27e3dc1cedc477a5f8939c

Observation b6c0dea8-bcd4-4dcd-8b03-81fbd12f337b · outbound

This paper cites TOFEDU: The Future of Education Journal, 3(5):1488–1499.

THiNK: Can Large Language Models Think-aloud? TOFEDU: The Future of Education Journal, 3(5):1488–1499

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.817646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:04:34.807017Z digest=sha256:4fd0bea737d64b6a314e411c3f7272e264bae5d3e1e0c2315c0b4fb29c103b03

Observation e335ab1c-82d5-46d0-9b99-a5dfdcb741e1 · outbound

This paper cites Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks.

THiNK: Can Large Language Models Think-aloud? Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.565123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.565123Z digest=sha256:95da776605c9e54d209879d152481afa00a8556c655b7a54e79f62eeb2082a7e

Pith citing papers

Observation 13673311-eda7-400a-b9d2-8c8c39141232 · inbound

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy cites this paper.

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy THiNK: Can Large Language Models Think-aloud?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:53.998837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:20:53.998837Z digest=sha256:9ad3cbbfb8821f9b22e7689fb3fa36ab6784375ca23f84b03a0b8435f4c251e5