Pith. sign in

Paper Citation Record · LEDGER

THiNK: Can Large Language Models Think-aloud?

As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2505.20184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20184 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:37.575936Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:53.998837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved15
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a7c8880-5fd9-469d-bcb4-c7d5e50c4e4d · outbound

This paper cites By analyzing which concepts are invoked, we assess the knowledge dimension activated during problem-solving and whether the LLM navigates these domains coherently.

THiNK: Can Large Language Models Think-aloud? By analyzing which concepts are invoked, we assess the knowledge dimension activated during problem-solving and whether the LLM navigates these domains coherently

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.254905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.403974Z digest=sha256:f6ec74abde951743648d99c4d1c881769c921dc753fdfd216d4aeb9307bd887e

Observation 0fa765ff-d851-481b-a821-0373e1944790 · outbound

This paper cites Thilo Hagendorff, Sarah Fabi, and Michal Kosinski.

THiNK: Can Large Language Models Think-aloud? Thilo Hagendorff, Sarah Fabi, and Michal Kosinski

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:44.055481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:34.621929Z digest=sha256:0b56999b86b07129f208ca372c84ccabe2dc8e8c59d5bc47049102930c72e7f9

Observation a4aa73a5-1796-43c5-aa1f-a76b5d6ed06c · outbound

This paper cites Representations are critical for logical coherence and traceability in reasoning.

THiNK: Can Large Language Models Think-aloud? Representations are critical for logical coherence and traceability in reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.892838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.570182Z digest=sha256:b995e9a258740c761556e54021aa51a7a230733136281ac8a41bd1fa872100ca

Observation 2f68695c-0270-47e6-81cb-67b468eebf67 · outbound

This paper cites A model’s ability to adapt its reasoning across such variants reflects generalization ability—an essential attribute of HOT.

THiNK: Can Large Language Models Think-aloud? A model’s ability to adapt its reasoning across such variants reflects generalization ability—an essential attribute of HOT

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.649080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.638056Z digest=sha256:202ba27b0e2014bfd7c796f07d0895397a64d45e5447061e2b6de67fb5f847c5

Observation d9d1a54d-e351-4c68-a515-5b41cb3329af · outbound

This paper cites Five Keys.

THiNK: Can Large Language Models Think-aloud? Five Keys

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.396321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.760718Z digest=sha256:e0ab165b1b3cd9f433db88917e23d8c4ff0c530c6b9f1632171d9fdbe6102302

Observation 30a63fb8-e0f4-41a6-a468-e77e64a0ac3f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

THiNK: Can Large Language Models Think-aloud? Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.989895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.989895Z digest=sha256:4059faf9c3cd70ada7d044f5b7f1b2b105823a1e3cdd385429bcb315339c2665

Observation 6f285043-ceae-4a7e-b7ac-a08237cb27f7 · outbound

This paper cites Assessing and Understanding Creativity in Large Language Models.

THiNK: Can Large Language Models Think-aloud? Assessing and Understanding Creativity in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:35.188891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:35.188891Z digest=sha256:17f6e5e0b10042e6a993d23565d1e2d58b9dc775ab932f22a24b7c8c1b2b745c

Observation eda686bc-321f-44f8-b9d2-085ef43b9912 · outbound

This paper cites Five Keys.

THiNK: Can Large Language Models Think-aloud? Five Keys

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.413353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.274686Z digest=sha256:96fc69978d0bb2b727daffbd94136f906a4bdbda15b2ae3d38731da4b4e80226

Observation fc6c6ce0-d87c-436f-b349-e6e86497d3cf · outbound

This paper cites These skills serve as proxies for prior knowledge and inform whether the LLM draws upon relevant background competence.

THiNK: Can Large Language Models Think-aloud? These skills serve as proxies for prior knowledge and inform whether the LLM draws upon relevant background competence

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.501537Z digest=sha256:820f58b18dc866e764018f8433c5d9634d2c5b5f90a1f7e57f354f67447050d0

Observation bc43f647-0c80-4f11-92ad-366e129e0133 · outbound

This paper cites You should try to understand and retrieve the specific mathematical information in it such as facts, patterns, objects, or contextual information, and decipher these meanings.

THiNK: Can Large Language Models Think-aloud? You should try to understand and retrieve the specific mathematical information in it such as facts, patterns, objects, or contextual information, and decipher these meanings

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:42.080713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.856691Z digest=sha256:afe96018655390c3e9fb7d7c3cf4fadd760ab85ab505382a7e5e020bbcf39144

Observation f80c2551-b0ae-441d-87f4-f27be7b9cd08 · outbound

This paper cites It includes understanding and organizing information, analyzing relationships, drawing conclusions, and distinguishing nuances.

THiNK: Can Large Language Models Think-aloud? It includes understanding and organizing information, analyzing relationships, drawing conclusions, and distinguishing nuances

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.709139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.926296Z digest=sha256:464e179f8e5f642d06b94ed4d23564bd945a345f6a74b04a423f16dc07b11dde

Observation 632569ac-4052-4187-86a6-dc67e329e041 · outbound

This paper cites These new expressions should have the same form as the given expressions in the previous generated math problem.

THiNK: Can Large Language Models Think-aloud? These new expressions should have the same form as the given expressions in the previous generated math problem

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.478200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:35.991143Z digest=sha256:2a5c260a31dc4fab609f62c2b6206d5a0b07ab2c961003e8deb86e6e83c6a27c

Observation d90fac4d-e73b-47fc-bc1f-49e61c612496 · outbound

This paper cites The generated stories must be a mathematical word problem with the corresponding expressions.

THiNK: Can Large Language Models Think-aloud? The generated stories must be a mathematical word problem with the corresponding expressions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:41.250540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.112437Z digest=sha256:dbec960ccc138ce2b1bc6a5ccbc0a43c75954d3965224034fc6074774cf19916

Observation 4ae966b1-69d3-4983-b64b-2753e21a22c2 · outbound

This paper cites question.

THiNK: Can Large Language Models Think-aloud? question

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:40.942421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.212697Z digest=sha256:de9f5ca881a0ee21fd9e606f1b1656be690a16339fe05605770d262a783e31a1

Observation d7221b4a-69f1-4292-9c7e-b07d16682968 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.731850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.299113Z digest=sha256:b5130ad30a26ebace6fc742e359c556a495cdd934b4082f8e530312fba776925

Observation 097e0f7a-e8c0-4aa7-8a53-1112802d3640 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.516278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.371403Z digest=sha256:5acf574f2f716abc312483c83df359133700856e4e7566559239be17ce568f32

Observation e7cca71a-fa57-4018-81a1-996ea4f8a584 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.346332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.467020Z digest=sha256:9f9c77b2426ac79c8fbdedd64f2cd7abb7971da2b456079d53d0d8bc141b52e9

Observation 71cfc3bd-c73a-4638-a83b-201870d4932c · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:40.119981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.557122Z digest=sha256:10288da32c871da8345826e7c8285eb366677f83737cdd6ef9cb4fccdb9cdbfb

Observation 0bd16435-849c-4b77-b423-35024e0a2a0d · outbound

This paper cites ID": null,.

THiNK: Can Large Language Models Think-aloud? ID": null,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:39.859244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.650966Z digest=sha256:32d99ff01d8d313bb8db2eb98da5af129b0d0d01d4537a502908ff6872709699

Observation 905ea987-80e9-4efc-ac45-90783e9ccf18 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.683488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.725952Z digest=sha256:28c0264ff18490ea8a96c7833bf07030f02683375b79198fbe0925a5d8621d76

Observation 373900a1-c4ed-414f-acab-352e4fb4840a · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.422480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.802182Z digest=sha256:30a998ee10cf38f251b7e1b5465d6fc264defd3ab9436fdde059514204aba6ab

Observation 28bbc1f8-ccd1-4e0f-a597-8346676859cb · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.228336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.871190Z digest=sha256:af0dc17af608aa28768be5670d7ea13e358d6add1b42f5561e7a030107e340d6

Observation 79f727b5-a1b4-4900-b207-266a76810f5d · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:39.013489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:36.981345Z digest=sha256:49c70002968e1ecc8e7b02efbf24641ba8bfe665bb6bf7d87348ad4b7ff42186

Observation fba5fb64-1ef4-447e-8f46-3851e00a2fc2 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:38.843275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:37.076085Z digest=sha256:aed6d74dddb1477d6ff22bf155fbd858ba2b6f3262f5a2facc20ecaffb33b188

Observation daaf974f-20f5-4e54-9b92-d6d3073a526b · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:38.604568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:37.196578Z digest=sha256:e077e792e6fc8f11618032ba6f144b63fc3a1032d85ea3feeb13040115e03e2f

Observation 88110793-a66b-414c-aaa9-dd9f02d9fef1 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 31

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.445619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:37.286441Z digest=sha256:c7b01caea9daffdbbd0b1d261fd76ec9175e79fee5578fc7f0e1614f6219c7f4

Observation acfc7e07-2c28-4479-9c14-bb5374b74b76 · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 32

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.191305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:37.367133Z digest=sha256:af406a79b4382b082cd8057f66c11854c12f2bc7889f6f4cd45aaecaf965ff2a

Observation 237518b0-cf69-410d-8cfa-93d6a34f5c0c · outbound

This paper cites an unresolved cited work.

THiNK: Can Large Language Models Think-aloud? Unresolved cited work

Reference 33

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:04:38.000547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:37.458862Z digest=sha256:ee01efb8b06c0f34054b39db9b0e4f6ddd071ee83353843b6a7eb88532fbd3af

Observation 7735d87f-a2ae-410b-84eb-b1424ad23de3 · outbound

This paper cites performance_score.

THiNK: Can Large Language Models Think-aloud? performance_score

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:37.798360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:37.575936Z digest=sha256:399acb8962e3afa4dd3f7e3cd7ce3a50a23b22ea429a3e2b0e054f9d226a7c84

Observation 6e3bc5a6-54ea-4d38-a129-9d44eba41060 · outbound

This paper cites English language teaching, 3(4):237– 248.

THiNK: Can Large Language Models Think-aloud? English language teaching, 3(4):237– 248

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.595243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:34.885086Z digest=sha256:caf1576291a1c7f9e1afccc2a85486a762649b71923564264ed4f3ef8dbdde11

Observation 4f4a0605-e26a-479d-bef5-6bda3f3c3127 · outbound

This paper cites MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models.

THiNK: Can Large Language Models Think-aloud? MulCogBench: A Multi-modal Cognitive Benchmark Dataset for Evaluating Chinese and English Computational Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:35.106825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:35.106825Z digest=sha256:5a3149c0bd3a18866920629ad4453c7dcea72126626f3edde8b691a85c164a7b

Observation 93bfd8cb-0986-4884-9228-76953960d739 · outbound

This paper cites HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models.

THiNK: Can Large Language Models Think-aloud? HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.719074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.719074Z digest=sha256:41558339b954b416c8bacbd29f558c824b05d59eed27e3dc1cedc477a5f8939c

Observation b6c0dea8-bcd4-4dcd-8b03-81fbd12f337b · outbound

This paper cites TOFEDU: The Future of Education Journal, 3(5):1488–1499.

THiNK: Can Large Language Models Think-aloud? TOFEDU: The Future of Education Journal, 3(5):1488–1499

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:43.817646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:34.807017Z digest=sha256:bac1e70f746b51c6aadd713319af61f360df87fdae81562cb1ca5e641774ef00

Observation e335ab1c-82d5-46d0-9b99-a5dfdcb741e1 · outbound

This paper cites Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks.

THiNK: Can Large Language Models Think-aloud? Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:34.565123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:34.565123Z digest=sha256:95da776605c9e54d209879d152481afa00a8556c655b7a54e79f62eeb2082a7e

Pith citing papers

Observation 13673311-eda7-400a-b9d2-8c8c39141232 · inbound

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy cites this paper.

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy THiNK: Can Large Language Models Think-aloud?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:53.998837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:20:53.998837Z digest=sha256:9ad3cbbfb8821f9b22e7689fb3fa36ab6784375ca23f84b03a0b8435f4c251e5