Pith. sign in

Paper Citation Record · LEDGER

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2506.17180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17180 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:31.684757Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c955ac35-2da0-4220-bc05-89b0955cb2b4 · outbound

This paper cites Assessment in science education: A study of teaching effectiveness.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Assessment in science education: A study of teaching effectiveness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.856619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:27.638495Z digest=sha256:4b14d4dfe91d2c2eefc5d188f17fbdfc08ff142f89a82ac052d93a47ff0b2f6b

Observation 6c8dde00-4a1f-49cc-b6f3-e5def3c6f121 · outbound

This paper cites Cbse assessment framework for science, maths and social science classes 9 and 10, 2020.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Cbse assessment framework for science, maths and social science classes 9 and 10, 2020

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.643752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:27.847326Z digest=sha256:ddfff29534753a29404aca702558122c8a8ca9818fc5256678b44de716a8bc10

Observation f70a2faf-7ce5-4f9d-9d97-80effefa7046 · outbound

This paper cites Bloom, Max D.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Bloom, Max D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.467416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:27.999453Z digest=sha256:8b63e3d180c59f5cc49f2a951b65fdec3ba94c174e2c993b926e2e7f807d50eb

Observation cc3aac27-7414-47c1-af97-732c5a7629cf · outbound

This paper cites Anderson, David R.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Anderson, David R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.304831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:28.216326Z digest=sha256:928c7a59a32b0a263b5ed2707093b43b25168f80c4223a894775294695c7c7d2

Observation 4e0c1e0d-557a-4186-b9a5-aec58b7c7245 · outbound

This paper cites Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:36.115224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:28.287009Z digest=sha256:342ec309c80de2a49a7f0c40290b76fcf5f7725b9321c7ab86186e0ad7cffc8b

Observation 5fa84c2e-fb64-4a47-ba99-f6ef48232ab1 · outbound

This paper cites Thinking, Fast and Slow.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Thinking, Fast and Slow

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:28.419985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:28.419985Z digest=sha256:66464804e54aec8f67810d25a4bad142afba9752fcef93eca670e3dbc6b8cc93

Observation 299d020f-5a79-4b5d-bbbf-7ddbf98535b0 · outbound

This paper cites Dual-process theories of higher cognition: Advancing the debate.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Dual-process theories of higher cognition: Advancing the debate

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.916792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:28.635181Z digest=sha256:d844b25f0ed438f92a0e9d5d15ee92017aeeb238056f30aef2b8144dc5e243ea

Observation ebf56e91-d8d5-4222-b94a-f513544a5419 · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.740210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:28.734129Z digest=sha256:5298d5a89cb6579dc9ae07e42f58190ae4708d80ef6342791c0c9f74dd2ffea6

Observation ddecb367-5cf1-4aaa-993a-984ad2228a55 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:35.621043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:28.877039Z digest=sha256:202ce0ac9cacf9b8aa1fb121ec989b4299a41614d414c336832a6d705efea9a4

Observation c2bccd31-ea7c-4933-9deb-3f9c479135be · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models SQuAD: 100,000+ questions for machine comprehension of text

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.456528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:29.049684Z digest=sha256:145deb9870fdddd4966569590aa8558c6bc0e4ce9d88292dfe3651bd31b96ade

Observation b4ec3bf0-a908-4e57-a880-2bf82157f0a2 · outbound

This paper cites Hutchinson, and Richard G.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Hutchinson, and Richard G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:35.200748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:29.223595Z digest=sha256:22d859936a7fad5ef1e0fb863898e3912c5cd643b939eddba703c474486c3c65

Observation 77217a69-f736-4e16-8949-2b0be8431aa2 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.960660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:29.379705Z digest=sha256:7393baeae8704cb568765cf06bf6848cd780b83a785186e4783a7c55c57cb08b

Observation 6c0ef4bf-850f-47b7-b6fc-291edba5b592 · outbound

This paper cites WinoGrande: An adversarial Winograd schema challenge at scale.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models WinoGrande: An adversarial Winograd schema challenge at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.641689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:29.555304Z digest=sha256:a0733960a7a6ce9fd24b24ef2fe15c1601c9e72ac0b2336a130184360368038d

Observation c78751fa-86a9-4360-a5a7-7734acbf13ae · outbound

This paper cites FEVER: A large-scale dataset for fact extraction and verification.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models FEVER: A large-scale dataset for fact extraction and verification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.476987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:29.641483Z digest=sha256:cf8aabdfd855faf350ef97ad7f4f8b5cd9d46f200afbf1b6cf74a6c80787e5ae

Observation 591bbfb5-3252-403c-af8e-d41d352e88e8 · outbound

This paper cites Fact or Fiction: Verifying Scientific Claims.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Fact or Fiction: Verifying Scientific Claims

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:29.816345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:29.816345Z digest=sha256:fe9c455e70f5e850defef6bb2ee4c5439e87150e36b42e210f3f070083a0a2a2

Observation 35b332ff-9fc1-4faa-aa6b-2531285d7808 · outbound

This paper cites The Llama 3 Herd of Models.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:29.937166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:29.937166Z digest=sha256:f6f62f45c9a4fe9af0b5d017b1d5ec03a5eeb3e7dec9db7990e96a4165fa70ad

Observation 32d3aeae-e6f9-4c04-9e40-fdcf7673f323 · outbound

This paper cites Qwen2.5 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Qwen2.5 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.032234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.032234Z digest=sha256:93d034842a64dc355f49fbf6f96b23bf041bcab7520b3c10dc2230834b708218

Observation f3837aa7-8513-464a-bf76-ca83285f287d · outbound

This paper cites Qwen3 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.125545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.125545Z digest=sha256:b939469fc03ebee0a0109061cb1d9bf4d9a5306251767e56061793d2a27fb863

Observation 96a14359-8ee3-4192-b5b0-ed03107b1a3a · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.266416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.266416Z digest=sha256:25a339fadafcf11ba156770c7f1298e814c0613993efc3bfd71b1d59c652baae

Observation ff585210-453b-490f-8e1d-f95af7e25606 · outbound

This paper cites Phi-4 Technical Report.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Phi-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:30.385458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:30.385458Z digest=sha256:7bd2f362614b24b24009c24a261c8eb40d8949daebe6f88b543629b661547920

Observation 6b1b6237-cd99-4231-83e4-e464c67bf3dd · outbound

This paper cites When one model casts doubt on another: A levels-of-analysis approach to causal discounting.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models When one model casts doubt on another: A levels-of-analysis approach to causal discounting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.297459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:30.554878Z digest=sha256:44927e1562afcbd122f3bf5b9e924b0bcb585e10c2231b72a22508f1a603b9e4

Observation 055c84a6-a36c-421e-ae31-884271148fc2 · outbound

This paper cites The most common error is incorrectly classifying option (b) as option (a) (176-214 instances across models).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models The most common error is incorrectly classifying option (b) as option (a) (176-214 instances across models)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:34.094387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:30.735614Z digest=sha256:af599183087b7e5c79fb3509913b7f3bbf4c7782759d8aed39e851835350b412

Observation 871c92f1-1083-405e-a5ba-60a540d572f7 · outbound

This paper cites Even the largest model, Qwen3-32B, misclassifies option (b) as option (a) in 195 cases (27.1% of all true (b) cases).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Even the largest model, Qwen3-32B, misclassifies option (b) as option (a) in 195 cases (27.1% of all true (b) cases)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:33.847413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:30.887249Z digest=sha256:8928906a3b95bc70bb4df770d0e910b878963912f997769538290a7882e1c90e

Observation fe6cdd58-1579-48e2-95e8-1f11e5c126c6 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:33.658462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.013302Z digest=sha256:9eafe062b9b3076b6ab90a1ea06d37c85c39b7aff3a702f78d58275e5facf1fd

Observation 4f8b2fdb-3e17-4b0c-99bd-2b73f6e7ca35 · outbound

This paper cites Model family differences suggest that some design and training strategies may better support causal reasoning, particularly in complex domains.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Model family differences suggest that some design and training strategies may better support causal reasoning, particularly in complex domains

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:33.339491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.082574Z digest=sha256:12d52569cb68cfe0a581d5abd17af284baf25a19af6f841ad2268559162cb3c9

Observation b4062b7d-825e-4ea3-841f-3f0893eead36 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:33.017115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.206215Z digest=sha256:3f7395a06423de003cabf1b6fd551961081993ed65212bddac0d95774aae5e23

Observation 47c042e4-5365-4b94-8326-4dc3c76dc8b1 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:32.804755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.284075Z digest=sha256:66c9c4d6a06a780b7e3d544ed12e5626dee0ef87ac4d88aa808b15b71b6daa2d

Observation d5df085c-b832-4137-8289-6784299fc9e8 · outbound

This paper cites These t-tests determine whether the differences in similarity are statistically significant or could have occurred by chance.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models These t-tests determine whether the differences in similarity are statistically significant or could have occurred by chance

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:32.563860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.431743Z digest=sha256:2269c503c3ccd2822589a16e58c9796ee8175d0eb1874ed58172268ebcc51220

Observation 58740287-4290-4ce3-8d25-b045a34a1198 · outbound

This paper cites Yes" versus all cases where it predicted.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Yes" versus all cases where it predicted

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:32.318389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.532676Z digest=sha256:285f9a0600cfc96e7c115f49bab5fbca76c60e5bbbaf3002193d1b7e91b30181

Observation 82369732-23bf-4c32-b2c9-944a32f473f5 · outbound

This paper cites an unresolved cited work.

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:32.123936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.612422Z digest=sha256:54243228577922224d1a8c3dbe95f11804360b20fc090328c51552bbc5488f5d

Observation b7dd928c-1127-47ae-bb7d-6a017bf98679 · outbound

This paper cites This pattern represents a fundamental confusion of correlation (semantic similarity) with causation (explanatory relationship).

CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models This pattern represents a fundamental confusion of correlation (semantic similarity) with causation (explanatory relationship)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:31.976303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:34:31.684757Z digest=sha256:d8879d429ea2e28ba8cb54ef87329e242619fefaea3ca73fa263318e75171f19

Pith citing papers

No inbound Pith citation observations are available.