Pith. sign in

Paper Citation Record · LEDGER

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy

As of 19 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2501.11721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11721 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:00:32.747694Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:36:55.345947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:36:55.445243Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9af8672c-e08a-41b2-91b7-49dbd3c8b608 · outbound

This paper cites Introducing gemini: Google's multimodal ai model.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Introducing gemini: Google's multimodal ai model

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.129677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.642092Z digest=sha256:8f0a068fd7c1b85ed862adc04fe9c98b6c23b6aa7fdde72d40cead55e722c822

Observation 604d3217-b67a-4864-83c9-53b324ee341a · outbound

This paper cites Introducing claude: Anthropic's ai assistant.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Introducing claude: Anthropic's ai assistant

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.114815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.647171Z digest=sha256:35536d53c48feb24f27fbd7c6d5c3ec3e45610deaccf1f3f9fc24abd5782a2d5

Observation 6c624207-06d6-40f1-b487-e8e488aa30e5 · outbound

This paper cites Explainability in ai: A survey.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Explainability in ai: A survey

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.100711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.651632Z digest=sha256:c4812717c6e5e799aea0f6e4da6aa25bfc8f5545ec6b8820c3047ad181e2e5a4

Observation 4745ce75-e64c-4df0-bb58-1a58cf4b7fb3 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 610--623, 2021.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 610--623, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.085857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.656481Z digest=sha256:0062914aa8b18cef27fc856af957712743e9e155de3d0177700de81375b29851

Observation 38fa1fa5-51b2-4ff2-9a2c-49782fe47314 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.660995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.660995Z digest=sha256:afd4dd48e4870d5f7d82537233229f4a439e86b55fa2e4b190a596cff7f9c680

Observation 3f885034-6a26-45aa-a436-1f61cc575181 · outbound

This paper cites Language models are few-shot learners.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Language models are few-shot learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.071485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.665972Z digest=sha256:ab4f6b54ab2a20617eca0dc3378d1ff3d588c6a0f09a6d750b78f0b089f8228c

Observation e7c0ecb4-fe86-414c-949e-8c707a9f56c5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.057548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.671018Z digest=sha256:f50f2ed120a04a54dc58631e190600523eb9c0b0c8c323ae060e8ff3e69ce68f

Observation 05283e8a-614a-4685-a41b-4bad9c1d6467 · outbound

This paper cites What if $\phi^4$ theory in 4 dimensions is non-trivial in the continuum?.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy What if $\phi^4$ theory in 4 dimensions is non-trivial in the continuum?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.675479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.675479Z digest=sha256:571f39400931bb4d241be82cb9c76a7783092eb136151fd41b4887901c15bfa1

Observation c537fa1b-b37e-464c-9e4f-0d61bb4d2692 · outbound

This paper cites Prediction of solar wind speed by applying convolutional neural network to potential field source surface (PFSS) magnetograms.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Prediction of solar wind speed by applying convolutional neural network to potential field source surface (PFSS) magnetograms

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T18:00:32.933485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.680162Z digest=sha256:1fcd6eaf696de72982ed5e9be295b74bd9332a768c3c331acd894c047458a2b2

Observation a725d6df-df7f-4458-add2-f214bc0435c1 · outbound

This paper cites Structure Factors for Hot Neutron Matter from Ab Initio Lattice Simulations with High-Fidelity Chiral Interactions.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Structure Factors for Hot Neutron Matter from Ab Initio Lattice Simulations with High-Fidelity Chiral Interactions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.684957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.684957Z digest=sha256:9d8a015d18a352d4f9edde8081d8e7dc554670be0f16fdbc1fb5f65759b00412

Observation d4f9bfda-7183-431a-8d62-7ff38e768035 · outbound

This paper cites XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:00:32.894951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.689053Z digest=sha256:8d4b4f2ccb67e32c9e01db1e5505b9145473b192bf7e9e0132a29beef2128889

Observation e7c513f4-769a-460a-86d0-0771b8c722e9 · outbound

This paper cites RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.693612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.693612Z digest=sha256:006101ac74b38d37a682e18beb0757509bf00283bef575bb6a7429c9c2de480e

Observation b06cca36-3cec-43ab-ba6d-052c9c216d86 · outbound

This paper cites Gpt-4 technical report.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Gpt-4 technical report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.697624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.697624Z digest=sha256:a198cb965118fac12560be0eecffa7aae2296ae2779b90612bafb58dbe6a1d19

Observation 29a10000-aa0b-4db0-ba6b-4d80c3c78123 · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Squad: 100,000+ questions for machine comprehension of text

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.034560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.701626Z digest=sha256:730626bb8ae851fea649b148bd9452aa9c4a21d4d792a726e5c756c23036e7a3

Observation 94671660-d958-470d-8465-82b5acbbfaf9 · outbound

This paper cites A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:00:32.852516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.705662Z digest=sha256:416432639bcce93fc357c39c0eaa4e1945b23ebea08dcba27029c73482177947

Observation 152ce31f-476b-490f-b0de-7a40c0496714 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.709790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.709790Z digest=sha256:e27b553549b5191d8ce1413d9890221fa739a66497479db956dfaed9f5554a5b

Observation 8cfcdcf8-cdd4-4a73-9b45-91dcf06e0348 · outbound

This paper cites On the fluid slip along a solid surface.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy On the fluid slip along a solid surface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.714286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.714286Z digest=sha256:4b9426576320f7c907586b42231c80ec85af41a750f31b570264488da1935e4d

Observation 3250ac9a-1429-4ea8-a333-a1313e7004d2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.718756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.718756Z digest=sha256:5182913a62fd7cdf2171277f8631ae80b58a03b5e389ce80c102f9c4ac70b447

Observation 14b953e3-a8be-47dc-998b-8f641d11d0b5 · outbound

This paper cites Teach me to explain: A review of machine learning interpretability through explanations.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Teach me to explain: A review of machine learning interpretability through explanations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.020554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.723277Z digest=sha256:d359dbc930faf83b6c5b4a13cb00b6c29977e55e6abd4a3e88c7bb7e21f85710

Observation 9cf92435-8d6a-446c-b62c-9041449ef5c9 · outbound

This paper cites Language Models can Evaluate Themselves via Probability Discrepancy.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Language Models can Evaluate Themselves via Probability Discrepancy

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.727925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.727925Z digest=sha256:38f1029f124cbd825afcb046ac71a4b8a8623a5694fb6fa332d3f18ccbac07e2

Observation aac56f11-00f1-4206-b11f-996977b5673c · outbound

This paper cites Evaluating paraphrase sensitivity in large language models.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Evaluating paraphrase sensitivity in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:00:33.006280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T18:00:32.732816Z digest=sha256:7a7e5fd9f3975e82b01d2c6ac16bf21cd40ff2959376225916e5447da5a9d116

Observation beac91c4-4df4-43ca-92f2-6504e68182b8 · outbound

This paper cites @esa (Ref.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy @esa (Ref

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.737661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.737661Z digest=sha256:d64c0448685f60745813f13cb76acc1d744662c210c037c0d1983d84758eea16

Observation ad4ab37a-ed28-4d22-bcd9-ffb4d37d84f9 · outbound

This paper cites an unresolved cited work.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.743000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.743000Z digest=sha256:ed8a57cd3a84ac3aec1e0b74607f106b48e15327b93ee0e0daca6d50c11b7a1b

Observation bded3b9a-c992-42de-ad0a-58f71e7b9702 · outbound

This paper cites an unresolved cited work.

Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:32.747694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:32.747694Z digest=sha256:b53f3abe5b64e8a188743bb82cbc2797b34b081b3476fb689d5f2bcc1a77b7de

Pith citing papers

Observation 4a41e5e2-253d-4009-bdaa-2fe060155d38 · inbound

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI cites this paper.

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:36:55.451186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:36:55.345947Z digest=sha256:0a5086dbbfa7abef4bf8094a3c018d7e3ade0093f7a22cbc7ff0f7d1f2858267