Pith. sign in

Paper Citation Record · LEDGER

Potemkin Understanding in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2506.21521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21521 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:30:32.913920Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:20:14.933677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:12:02.502599Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d63f77aa-6ef6-49ee-a283-79eb99afab8a · outbound

This paper cites write newline.

Potemkin Understanding in Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.435106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.435106Z digest=sha256:f466656ba7724f23684f4b587bb1aeda34436bc69cd1c963fc6d5b9a3ee6d6fe

Observation f1e6da37-5ba4-4f69-8fa9-0c59c9d581ff · outbound

This paper cites GPT-4 Can't Reason.

Potemkin Understanding in Large Language Models GPT-4 Can't Reason

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.514523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.514523Z digest=sha256:ff2e12b7d15f1cca9e45e5749043f18a88d3bc6ac7277c4cd2fb62587049f903

Observation f2924307-bf4f-46d0-8620-46c553022461 · outbound

This paper cites Synthetic and Natural Noise Both Break Neural Machine Translation.

Potemkin Understanding in Large Language Models Synthetic and Natural Noise Both Break Neural Machine Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.626627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.626627Z digest=sha256:a166d332b2bac462499b98d90b2bc351d258a68a176b53fff8ad7f8ac6e28bd3

Observation da117aa1-2c46-4bef-ab57-ea86758a63f7 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Potemkin Understanding in Large Language Models The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.715765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.715765Z digest=sha256:e8a5a2c262e980f19a9636c8ec7dc4339be8952388e6406e8726e99d60c061f4

Observation 2aadb985-3932-4443-9b5c-20ebf50d25b2 · outbound

This paper cites What Will it Take to Fix Benchmarking in Natural Language Understanding?.

Potemkin Understanding in Large Language Models What Will it Take to Fix Benchmarking in Natural Language Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.809315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.809315Z digest=sha256:b1c56a3df73bb21fa1b56dae33f89eed08fb3530d0b832bb5e413f97b014f28f

Observation acad158d-780c-4e01-9d25-2f7b9e40fe4f · outbound

This paper cites A large annotated corpus for learning natural language inference.

Potemkin Understanding in Large Language Models A large annotated corpus for learning natural language inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.890401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.890401Z digest=sha256:a0e33cd99577d6809b30473f41d42abbd4ef6b4c51c13c63fdc3b214e59be42c

Observation 2b23d7a8-3d0b-4c30-9521-1a1bfc1a288f · outbound

This paper cites T., Li, Y., Lundberg, S., et al.

Potemkin Understanding in Large Language Models T., Li, Y., Lundberg, S., et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:37.175863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:28.992268Z digest=sha256:74af055274989ef948cdd990006995bdd1a3dc98b560d73406c3508d426ed457

Observation 3c13f625-6a91-4d56-bd70-58eeb0d90de1 · outbound

This paper cites With Little Power Comes Great Responsibility.

Potemkin Understanding in Large Language Models With Little Power Comes Great Responsibility

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:30:33.752650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.088514Z digest=sha256:1a5d0659882634ebe13627ad0032221f4c0a4cce3e85f8025af835d118b435aa

Observation 570370e0-e76e-4430-9284-c1068694b47e · outbound

This paper cites ChatBench: From Static Benchmarks to Human-AI Evaluation.

Potemkin Understanding in Large Language Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:29.181666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:29.181666Z digest=sha256:6d5eba1d354b86c3d0a0c961d294d0832c1df341326833b053170280d2452015

Observation ebe472ad-7448-4859-924d-4375114536cb · outbound

This paper cites N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J.

Potemkin Understanding in Large Language Models N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:37.015281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.278520Z digest=sha256:f89de76500be8f3bc708725d21bc8b8f999ede833b4de1e915d446456edfe1d9

Observation f0562646-d1d9-4fd3-b47a-e047e37a74de · outbound

This paper cites an unresolved cited work.

Potemkin Understanding in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:30:36.872633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.378079Z digest=sha256:fb853360198121d0013d94df0870f916298100ebb86ce13cce5329786beaed5e

Observation c9b3191d-80f6-4b06-9503-beaa37bb8d3c · outbound

This paper cites and Etzioni, O.

Potemkin Understanding in Large Language Models and Etzioni, O

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.720223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.481957Z digest=sha256:79d2ecfca8866ef71d6f225965b69d2188e044aeb066b1d32b531f560a050873

Observation 1ff9d6c6-be20-479b-8761-9cabe6a83991 · outbound

This paper cites Recognizing textual entailment: Rational, evaluation and approaches--erratum.

Potemkin Understanding in Large Language Models Recognizing textual entailment: Rational, evaluation and approaches--erratum

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.600044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.559562Z digest=sha256:9ceea3216acb6c8c3217918a8dee4a4edb9ad4e11300b40f00af002ea4066ef3

Observation 4cb16b18-fd89-4569-94ca-464bf44350d6 · outbound

This paper cites Testing ai on language comprehension tasks reveals insensitivity to underlying meaning.

Potemkin Understanding in Large Language Models Testing ai on language comprehension tasks reveals insensitivity to underlying meaning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.453392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.621688Z digest=sha256:c81bf1f9d4678342d80a90b53eee9d70759efa47b449ea27a409f3107bfc1e81

Observation a0ea739a-535a-4172-8f46-2181181d8d19 · outbound

This paper cites and Meurers, D.

Potemkin Understanding in Large Language Models and Meurers, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.269650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.666182Z digest=sha256:b8ad0e782ab924002f63b39ac10a9031b29fec4edc0764783ca501eb442bba26

Observation 6077b6e3-a40e-4ede-a49f-cc31384401c8 · outbound

This paper cites Measuring and improving consistency in pretrained language models.

Potemkin Understanding in Large Language Models Measuring and improving consistency in pretrained language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.131723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.750083Z digest=sha256:a4056290dc242c9d6b2ee3c6eb07e631dcbb708145461f4d7516caba6d4c903d

Observation 795b1987-0c86-46cd-9f60-e2fbbe755337 · outbound

This paper cites Evaluating superhuman models with consistency checks.

Potemkin Understanding in Large Language Models Evaluating superhuman models with consistency checks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.945909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.825513Z digest=sha256:f85059795a823c541f6f0fa08c09d780346e283bb13f2f0ccb10a7482680104c

Observation df50e5bb-3e3c-4b1c-88ec-df9aaa6f1fe7 · outbound

This paper cites W., Wallach, H., Iii, H.

Potemkin Understanding in Large Language Models W., Wallach, H., Iii, H

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.807415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.888225Z digest=sha256:9991b38d3081b15c9cf027a4444ffe72ac1e9826a41eb35bb9bd6b3b95ebe846

Observation 4ef75983-0583-4e5c-bdb5-f556c86df559 · outbound

This paper cites an unresolved cited work.

Potemkin Understanding in Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:30:35.689875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.949572Z digest=sha256:a8fb7a17dfab276a00b1dd910c9e3ef3c72011f63b9f90c0713b6741ba9457bf

Observation 9d7a95cd-1d8f-4e3a-ad3e-c7466b06c8a0 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Potemkin Understanding in Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.028454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.028454Z digest=sha256:20527260c8bdcd620d6d1cb58c9f0eb7d712b337605e0e693d0e14eaa1ff4ab5

Observation ad7f7eb0-fd01-4e0f-82f0-18976d86aa83 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Potemkin Understanding in Large Language Models Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.090564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.090564Z digest=sha256:7400a941bb22d6edfcc6c4512abe0c36d6a7b308e0630d9bcfdca6fa0e092450

Observation 882665b5-e390-4393-a700-f7ee81ed8000 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Potemkin Understanding in Large Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.137112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.137112Z digest=sha256:d7908f6005fc2c7622989e4fbef550e36e58d9c69940d558c50326f6fe9c3817

Observation 57ecf5a2-459b-4cc0-8b2b-0afe2ce9ee84 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

Potemkin Understanding in Large Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.569299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.184234Z digest=sha256:e0603a20492ec8a8abf2859932eaa3e9c6d7b117df36d7668b7e834b64029f7c

Observation bb395e0f-87b7-47c5-b22f-19332a00494b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Potemkin Understanding in Large Language Models Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.241460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.241460Z digest=sha256:8b30962e22bccf17a248d9a1b8c36bcd0b8cbe9b2237aaf46bbbc951ce9d69d5

Observation 9ed4d6bd-64af-43c2-8e0e-b91ee427f5ac · outbound

This paper cites Understanding by Understanding Not: Modeling Negation in Language Models.

Potemkin Understanding in Large Language Models Understanding by Understanding Not: Modeling Negation in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.327528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.327528Z digest=sha256:3467b395084f538dc94fa40fe05abb4a8e31845aa6386f3f07824a389b12b78b

Observation f80f7806-56f6-4f3a-adde-ab4af24c94be · outbound

This paper cites Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models.

Potemkin Understanding in Large Language Models Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.380612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.380612Z digest=sha256:5e39b3688f50defa67351c7d39f0c2f32a051cea23f15ebe35895723e184418f

Observation 0a12bff8-1508-4e41-9956-e1d2c96a5a6b · outbound

This paper cites Adversarial Example Generation with Syntactically Controlled Paraphrase Networks.

Potemkin Understanding in Large Language Models Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.427899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.427899Z digest=sha256:fe4dae1602af19eb9dc3bc027bef3c34c92f092b7d8fc932bcd67e3e1372a5c2

Observation bd4c4a24-c70e-48cf-a109-fd0779d6db79 · outbound

This paper cites Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models.

Potemkin Understanding in Large Language Models Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:30:33.511891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.469419Z digest=sha256:ad358f78725d872415fed960e8baca7f623741105d10ab11d75032fa04e7aec7

Observation 4ecb5384-98e1-4909-99a9-00cfd77fb317 · outbound

This paper cites S., and Lukasiewicz, T.

Potemkin Understanding in Large Language Models S., and Lukasiewicz, T

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.468299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.557719Z digest=sha256:e894c87e8ddf3c9efcdd18a2cf2969b12259c3e07cedc4b6e37cd598d5fcaff1

Observation 68983c8b-c899-4210-9d69-74e87fd3f5a1 · outbound

This paper cites Consistency Analysis of ChatGPT.

Potemkin Understanding in Large Language Models Consistency Analysis of ChatGPT

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.615032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.615032Z digest=sha256:4c6c8f311539c45477624386c33dc4ca7443070bca9264f11a368bf650cc99f7

Observation 354c96c6-6b8a-4b97-9c6b-027e9347c2bc · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

Potemkin Understanding in Large Language Models What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.375868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.655859Z digest=sha256:53e502b3655d064498a2f2201048c5857ea46e0f480c2829e975c5097628e90c

Observation 6c072be9-1d61-4f5e-97e5-5ac8b700aab8 · outbound

This paper cites Dynabench: Rethinking Benchmarking in NLP.

Potemkin Understanding in Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.708684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.708684Z digest=sha256:880c835307e66321b76a55a39bdc43f12c2e9024ed3cbb3c33299d3a8cc50e75

Observation 0186b315-2370-42f2-bcec-d4c750206225 · outbound

This paper cites a ldchen, S., Binder, A., Montavon, G., Samek, W., and M \.

Potemkin Understanding in Large Language Models a ldchen, S., Binder, A., Montavon, G., Samek, W., and M \

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.273866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.785634Z digest=sha256:2dbf91e810f651b120d0b5050e988ab1561622f44addefdd701b630fa3e1d621

Observation 61a4b5d1-70c0-452f-accd-f2ac542cbe02 · outbound

This paper cites Benchmarking and Improving Generator-Validator Consistency of Language Models.

Potemkin Understanding in Large Language Models Benchmarking and Improving Generator-Validator Consistency of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.853855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.853855Z digest=sha256:29b87ae0d34959d0292bcf4d208c0cdeed556f54680576752cb5b79f7def2c6d

Observation 34dd9cdc-8b85-4a39-9792-2ecfad2c4309 · outbound

This paper cites Holistic Evaluation of Language Models.

Potemkin Understanding in Large Language Models Holistic Evaluation of Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.910951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.910951Z digest=sha256:674d39f65b9c916e2be805e740f3c62a876203c520571a594e229523c322276b

Observation 21d94355-3964-424e-89e1-753c74677884 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Potemkin Understanding in Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.965179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.965179Z digest=sha256:25f1e8a9b5518fcfc6401bd0778403310ca17624d28c5aef305b0e1508bd4925

Observation 2549b930-ffc8-4b3d-934c-a762e4745d65 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

Potemkin Understanding in Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.067924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.067924Z digest=sha256:2b1ebe4419e232245101b391f6618da9d5804eff65a127b3341ba82bcf5fa1dd

Observation be411334-3402-45b3-8c6d-ce53abf7f192 · outbound

This paper cites Sycophancy in Large Language Models: Causes and Mitigations.

Potemkin Understanding in Large Language Models Sycophancy in Large Language Models: Causes and Mitigations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.121504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.121504Z digest=sha256:12bb97c6bdc63261efba079176b1c2d68cdb730ff6759d405e197c61dcee564a

Observation 091c2206-b1ef-4a53-bb76-a814cce90897 · outbound

This paper cites Locating and editing factual associations in gpt.

Potemkin Understanding in Large Language Models Locating and editing factual associations in gpt

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.195017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.176258Z digest=sha256:10ebea836529856f1db3cb123d54cae525f450ecec28af079e17f05b6713e6db

Observation 6035467f-3a81-400a-b0d5-cc888ce0ad89 · outbound

This paper cites Fast Model Editing at Scale.

Potemkin Understanding in Large Language Models Fast Model Editing at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.224912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.224912Z digest=sha256:931a645ac786f78de74418738efab5c8f7a980f418bba67c9567861b8063400f

Observation f5d35c09-92f8-4221-b27e-3e9898934714 · outbound

This paper cites Why AI is Harder Than We Think.

Potemkin Understanding in Large Language Models Why AI is Harder Than We Think

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.291743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.291743Z digest=sha256:6ec48110acc07c4b9528ebc9fc233159f8972786b462fc639bd9bd4dc1541bf7

Observation 3f0bdea2-db8d-4924-b014-b67fb3794e51 · outbound

This paper cites D., Bender, E.

Potemkin Understanding in Large Language Models D., Bender, E

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.113972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.379058Z digest=sha256:6de3918ab25e6459ab737209baa2383a49c02f8712374843c8b04b39a7aa511d

Observation 4d8bd2d3-03dd-41a4-b40a-9090106f3e7f · outbound

This paper cites Language Models as Knowledge Bases?.

Potemkin Understanding in Large Language Models Language Models as Knowledge Bases?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.430923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.430923Z digest=sha256:8b5eeac3f2c5bfbec6638eaa4ca0cd985eb3b6c75790c8fa0422eddfa556bc9b

Observation c7fccb6c-5ef1-4c84-ad1c-c551ef3259a2 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Potemkin Understanding in Large Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.483348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.483348Z digest=sha256:d135acb7fc2db9fd1d0b4435db0cf40fef831e992b63a84f279f82edcea2792f

Observation 6c20d737-f59a-4e24-a439-fe41835ade21 · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

Potemkin Understanding in Large Language Models AI and the Everything in the Whole Wide World Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.598191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.598191Z digest=sha256:dfbca0eb080ba4d245cf4fd7ffd3a79f16b1cf2ca8a94d1f9d5084c52267d511

Observation 8654d9be-99c3-404f-9f7e-8492fb9d6de5 · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International conference on machine learning.

Potemkin Understanding in Large Language Models Do imagenet classifiers generalize to imagenet? In International conference on machine learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.017358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.642476Z digest=sha256:08cb92faee552d5fad98aea84ee96f920480a009336ee110f094032364e16c20

Observation efe6e90d-51f8-4d07-8a1b-e0d52e412b8e · outbound

This paper cites BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.

Potemkin Understanding in Large Language Models BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.722951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.722951Z digest=sha256:c3844304a44b9c5a91f761441b5d536430a994afbbd36fafaef2fcb8b3c4f2b0

Observation bf318aeb-57fe-4762-ba40-f4bb1662706d · outbound

This paper cites why should i trust you?.

Potemkin Understanding in Large Language Models why should i trust you?

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.941864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.791629Z digest=sha256:a53f9e6d5df54ef8afb0c463af1b1f14be8c4f8af4b1b3d4f50c60015a7feee3

Observation 59283c74-da0b-44ab-9b13-74f0e61b7f41 · outbound

This paper cites T., Singh, S., and Guestrin, C.

Potemkin Understanding in Large Language Models T., Singh, S., and Guestrin, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.829170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.837348Z digest=sha256:fb8f656c0580e01086c0d82a203b0f8f25ecbb0c7f3edc8c01a9245736b8b834

Observation d13db075-ea71-430e-a719-e1007ab4d794 · outbound

This paper cites T., Guestrin, C., and Singh, S.

Potemkin Understanding in Large Language Models T., Guestrin, C., and Singh, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.720084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.909318Z digest=sha256:e75004d115d01de11f064fd92b8e52affa2cc995f812c37115c8a4ed604c6478

Observation 1bbb4541-0b04-4f46-83a6-dcf80abbd733 · outbound

This paper cites Beyond Accuracy: Behavioral Testing of NLP models with CheckList.

Potemkin Understanding in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.993054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.993054Z digest=sha256:ca6f87c8dba615431098904e8612aee7a28444c2fb61d448aefdefcaa82fbdd6

Observation 3e4ec196-89d8-4a30-a34d-ea2f8cffbdc8 · outbound

This paper cites Models in the wild: On corruption robustness of neural nlp systems.

Potemkin Understanding in Large Language Models Models in the wild: On corruption robustness of neural nlp systems

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.619733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.060512Z digest=sha256:d4c08e14d335b5fa393533fd245f72aef199befc21b69afc88bd4d67c22725ee

Observation 756778bd-8b0b-448d-86bc-517d7c6e4844 · outbound

This paper cites LLMs' Understanding of Natural Language Revealed.

Potemkin Understanding in Large Language Models LLMs' Understanding of Natural Language Revealed

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:30:33.173737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.137478Z digest=sha256:6a8e7b096de575f64eae25218973206ef1d0d8be0498b184e5efc612a3b06649

Observation 2cc69fd1-3dd3-41b0-864e-6fd6c07ae10a · outbound

This paper cites everyone wants to do the model work, not the data work.

Potemkin Understanding in Large Language Models everyone wants to do the model work, not the data work

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.533005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.192447Z digest=sha256:89aa8ae22be830d765707ce0af556e8d6744b760825c9f13b257da5a46ee0925

Observation 216fad6c-c506-449f-8b23-6b8a3e88fbca · outbound

This paper cites H., Sch \"a rli, N., and Zhou, D.

Potemkin Understanding in Large Language Models H., Sch \"a rli, N., and Zhou, D

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.449212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.275321Z digest=sha256:cfcf2a804f5c06c0f9f37c1e8f229219559d4b70bcdfac47a0c32e7a4131c807

Observation 016182a3-c9ea-4dc4-b952-cf87cc9e0a17 · outbound

This paper cites and Choi, Y.

Potemkin Understanding in Large Language Models and Choi, Y

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.348695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.351314Z digest=sha256:82ec8feb8803c7f61eaf9eaabaa606f74f9dd1cc34d5bd616e9b415b31457da8

Observation 14bf4e05-0d71-4d7b-82bf-cfe53b9e85b6 · outbound

This paper cites S., Wei, J., Chung, H.

Potemkin Understanding in Large Language Models S., Wei, J., Chung, H

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.232060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.415105Z digest=sha256:1eb22d3e669f384670d4c4dff782617bdc14cfe39f4c085b85834732e4ecedb1

Observation fc60529a-e377-4d7f-bff1-68ebb3aa1c42 · outbound

This paper cites Evaluating the Factual Consistency of Large Language Models Through News Summarization.

Potemkin Understanding in Large Language Models Evaluating the Factual Consistency of Large Language Models Through News Summarization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.462098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.462098Z digest=sha256:7d6ad23950c7ad8ed7a56b98c27d31e87fe7372b9d0a981e738981f3278d795d

Observation e7ad8acb-9e25-4e00-a2f2-9ec75fa9b2f0 · outbound

This paper cites Evaluating the World Model Implicit in a Generative Model.

Potemkin Understanding in Large Language Models Evaluating the World Model Implicit in a Generative Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.551668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.551668Z digest=sha256:13fb378c15ccd8e289ac533f918884f739e1f430103fc708b9c237bb58fad1b5

Observation c9c5fe9e-7c31-43b3-91c9-c7e8ec64e601 · outbound

This paper cites Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function.

Potemkin Understanding in Large Language Models Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.594808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.594808Z digest=sha256:dee32d05333f2dc2a9035b4422eb43191df44907a6e5131d90cb96f948ba40c3

Observation c157b09a-1b97-49fd-b6bb-facc38aef7bc · outbound

This paper cites On the planning abilities of large language models-a critical investigation.

Potemkin Understanding in Large Language Models On the planning abilities of large language models-a critical investigation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.145704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.648173Z digest=sha256:ef27dd122ba048c3e38bcd61515923fbba105262788af1cfe49d08e47f89a11f

Observation ee7bcbd8-a32b-439c-88b7-95a1bbc9be8c · outbound

This paper cites T., Heer, J., and Weld, D.

Potemkin Understanding in Large Language Models T., Heer, J., and Weld, D

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.056585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.704793Z digest=sha256:2d35f70b4e59aeb5f30ca872659e5a163b2ef7326c6aaee3bd6f0a7a8c6bd650

Observation c4931cb1-6d7c-4e55-a8b2-411fe52458cc · outbound

This paper cites Kformer: Knowledge injection in transformer feed-forward layers.

Potemkin Understanding in Large Language Models Kformer: Knowledge injection in transformer feed-forward layers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:33.948182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.766649Z digest=sha256:eacfe3c0eddea54de24ba32c4b4d4619d943083bf7a406f807fb084788b496f8

Observation 360749ae-3ffb-43ba-a9d3-9645dd258d68 · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Potemkin Understanding in Large Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.834986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.834986Z digest=sha256:b3657e1712034de850503848c3f8c8a105288633a44ec926a7e041cb1fedb8dc

Observation 0468593e-4a93-42a7-8b83-bf4510fa936e · outbound

This paper cites Modifying Memories in Transformer Models.

Potemkin Understanding in Large Language Models Modifying Memories in Transformer Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.913920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.913920Z digest=sha256:dc2b57676565d674dc9046e84d06f293b226fcb211f9a56430c65257d22d47bb

Pith citing papers

Observation e47380ed-7f27-40ba-9c70-a07012ce052f · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Potemkin Understanding in Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.504578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:4e0fc0dd40c17cfb6b640733300db24944bf6639e2ae6b75d50bb338c034540f

Observation b920a734-a6e2-4c40-8dfb-4b2310026e8f · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny Potemkin Understanding in Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:14.933677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:14.933677Z digest=sha256:c2086bc1979b5221f83080752a609b6098c50d745ffb0232f2e7b096b9417315

Observation d3377b6c-ba71-43ca-a1e9-e72554f74392 · inbound

The wall confronting large language models cites this paper.

The wall confronting large language models Potemkin Understanding in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:12:02.086965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:12:02.086965Z digest=sha256:f3ad3a5428fc98c1e34b487d2d37940a2d4116cd069663f1b3252586bb032902

Observation 4f151f5c-19ea-4c68-a39f-a7fdbbbd2f34 · inbound

A paradox of AI fluency cites this paper.

A paradox of AI fluency Potemkin Understanding in Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:52.218393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T16:17:48.531790Z digest=sha256:4841cde36c68f92cc5d562d668a6eea4efab7c5d6f74c07ae99a4293f8f44000

Observation 23d6e3c3-18ed-4570-9836-ee238b71a00a · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Potemkin Understanding in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:edf8517827776b72dc8c838aec9292692bb115d58294b2efc0e290d621bb4fe1