Pith. sign in

Paper Citation Record · LEDGER

Potemkin Understanding in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2506.21521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21521 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:30:32.913920Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:20:14.933677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:12:02.502599Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d63f77aa-6ef6-49ee-a283-79eb99afab8a · outbound

This paper cites write newline.

Potemkin Understanding in Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.435106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.435106Z digest=sha256:b10a11baca49eb34d8946e4ef3b3fc551679ba472183d442433f7acb6454f2b1

Observation f1e6da37-5ba4-4f69-8fa9-0c59c9d581ff · outbound

This paper cites GPT-4 Can't Reason.

Potemkin Understanding in Large Language Models GPT-4 Can't Reason

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.514523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.514523Z digest=sha256:6acbf69e96d9455174eb31db16eb35fba913cb819c10bc3f94e8a4220f03b300

Observation f2924307-bf4f-46d0-8620-46c553022461 · outbound

This paper cites Synthetic and Natural Noise Both Break Neural Machine Translation.

Potemkin Understanding in Large Language Models Synthetic and Natural Noise Both Break Neural Machine Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.626627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.626627Z digest=sha256:c9d960c8ca4ff41844494ae1b137dbaee80ebae6066a4d7845a532b64977ec55

Observation da117aa1-2c46-4bef-ab57-ea86758a63f7 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Potemkin Understanding in Large Language Models The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.715765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.715765Z digest=sha256:435213a181abf1038ef436a75618e9d81020fdafbb2dd6ed35ff01581c5aca80

Observation 2aadb985-3932-4443-9b5c-20ebf50d25b2 · outbound

This paper cites What Will it Take to Fix Benchmarking in Natural Language Understanding?.

Potemkin Understanding in Large Language Models What Will it Take to Fix Benchmarking in Natural Language Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.809315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.809315Z digest=sha256:8645f281b3e965004287504ef968afeda221350f379d0ff0e4baf769c474cc85

Observation acad158d-780c-4e01-9d25-2f7b9e40fe4f · outbound

This paper cites A large annotated corpus for learning natural language inference.

Potemkin Understanding in Large Language Models A large annotated corpus for learning natural language inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:28.890401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:28.890401Z digest=sha256:b271c4354afd1ff8dc9006903f0a6d16b880d2a5cab68bacec12b259617a616c

Observation 2b23d7a8-3d0b-4c30-9521-1a1bfc1a288f · outbound

This paper cites T., Li, Y., Lundberg, S., et al.

Potemkin Understanding in Large Language Models T., Li, Y., Lundberg, S., et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:37.175863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:28.992268Z digest=sha256:f7da4b38a0f7d1980d18f7665e520c83add3a0eb63fc97e2bdb7068e80e632e5

Observation 3c13f625-6a91-4d56-bd70-58eeb0d90de1 · outbound

This paper cites With Little Power Comes Great Responsibility.

Potemkin Understanding in Large Language Models With Little Power Comes Great Responsibility

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:30:33.752650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.088514Z digest=sha256:a9ef33d1a2b0b524466ae253a223bc1328278b58d718932b89f7d1a5c39a3c32

Observation 570370e0-e76e-4430-9284-c1068694b47e · outbound

This paper cites ChatBench: From Static Benchmarks to Human-AI Evaluation.

Potemkin Understanding in Large Language Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:29.181666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:29.181666Z digest=sha256:38c473fbc61e06aa2de1de25bd17e27c2311e6579a9749ee29192c593ba4892c

Observation ebe472ad-7448-4859-924d-4375114536cb · outbound

This paper cites N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J.

Potemkin Understanding in Large Language Models N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:37.015281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.278520Z digest=sha256:6347acfc56fba9b141c4009ee9149bfd9f528c0813d4aa7d924c3706ecb281db

Observation f0562646-d1d9-4fd3-b47a-e047e37a74de · outbound

This paper cites an unresolved cited work.

Potemkin Understanding in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:30:36.872633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.378079Z digest=sha256:dcd432b281c8dd89a5f2c1e2184ab5122d85f0416a5c08829926f2bcea9031d9

Observation c9b3191d-80f6-4b06-9503-beaa37bb8d3c · outbound

This paper cites and Etzioni, O.

Potemkin Understanding in Large Language Models and Etzioni, O

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.720223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.481957Z digest=sha256:3407a3cab87aff16bc2459b97db81283e65ede98c07e086661d4ba769813122c

Observation 1ff9d6c6-be20-479b-8761-9cabe6a83991 · outbound

This paper cites Recognizing textual entailment: Rational, evaluation and approaches--erratum.

Potemkin Understanding in Large Language Models Recognizing textual entailment: Rational, evaluation and approaches--erratum

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.600044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.559562Z digest=sha256:03af5be8339d72d64c88c49969079d9f371ed057e7d85751e8b55b898efb2c79

Observation 4cb16b18-fd89-4569-94ca-464bf44350d6 · outbound

This paper cites Testing ai on language comprehension tasks reveals insensitivity to underlying meaning.

Potemkin Understanding in Large Language Models Testing ai on language comprehension tasks reveals insensitivity to underlying meaning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.453392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.621688Z digest=sha256:94faf1f95a03a9f53684605231f247b00ffd242ae1a6ab8fcb8beffad4893b48

Observation a0ea739a-535a-4172-8f46-2181181d8d19 · outbound

This paper cites and Meurers, D.

Potemkin Understanding in Large Language Models and Meurers, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.269650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.666182Z digest=sha256:71a4a12ea952999c4ba3368eba481e40f89586ed5c62c3fd527965ac0670bf02

Observation 6077b6e3-a40e-4ede-a49f-cc31384401c8 · outbound

This paper cites Measuring and improving consistency in pretrained language models.

Potemkin Understanding in Large Language Models Measuring and improving consistency in pretrained language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:36.131723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.750083Z digest=sha256:2cd58f258c78750a379d6cd4a5ab60da56163a88d3fdee70931cf59346c20023

Observation 795b1987-0c86-46cd-9f60-e2fbbe755337 · outbound

This paper cites Evaluating superhuman models with consistency checks.

Potemkin Understanding in Large Language Models Evaluating superhuman models with consistency checks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.945909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.825513Z digest=sha256:7c7f05fbd75ceec5dca6f00bd83d1120080fdea0931c888d7e3ba6c8afdfc192

Observation df50e5bb-3e3c-4b1c-88ec-df9aaa6f1fe7 · outbound

This paper cites W., Wallach, H., Iii, H.

Potemkin Understanding in Large Language Models W., Wallach, H., Iii, H

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.807415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.888225Z digest=sha256:559718c1ab91b07dbbeb5f71eb4d18648bbc12202129f90a9eb93e75ed442052

Observation 4ef75983-0583-4e5c-bdb5-f556c86df559 · outbound

This paper cites an unresolved cited work.

Potemkin Understanding in Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:30:35.689875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:29.949572Z digest=sha256:e53f47e2ef3bee401097a65d7e0193ee412de4d5787f54d841bedf16a8838e26

Observation 9d7a95cd-1d8f-4e3a-ad3e-c7466b06c8a0 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Potemkin Understanding in Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.028454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.028454Z digest=sha256:ff002daf76f60b196ed6c03ecd9b33c1d4eff2076fc2b15e9b3b049f1f229a59

Observation ad7f7eb0-fd01-4e0f-82f0-18976d86aa83 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Potemkin Understanding in Large Language Models Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.090564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.090564Z digest=sha256:4b407cd532364a820c5f50b503a28a8f3776614c6df52fb44733b8153328e0b2

Observation 882665b5-e390-4393-a700-f7ee81ed8000 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Potemkin Understanding in Large Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.137112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.137112Z digest=sha256:44074d5abb038ee454799625e76723006b09b3d432728da989c5516b910da4a6

Observation 57ecf5a2-459b-4cc0-8b2b-0afe2ce9ee84 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

Potemkin Understanding in Large Language Models What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.569299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.184234Z digest=sha256:83cc67a857e2854ab70fe91b2fa9b046ec085799dd0cf7fe89cec4132c6ddf36

Observation bb395e0f-87b7-47c5-b22f-19332a00494b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Potemkin Understanding in Large Language Models Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.241460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.241460Z digest=sha256:c0744832bef95b2d3b326c91615b247a205d736b6cef3caa663a6b4d782eac16

Observation 9ed4d6bd-64af-43c2-8e0e-b91ee427f5ac · outbound

This paper cites Understanding by Understanding Not: Modeling Negation in Language Models.

Potemkin Understanding in Large Language Models Understanding by Understanding Not: Modeling Negation in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.327528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.327528Z digest=sha256:17e9dce07bf9c6d233471c3bea5fc9e65fae329a2de433b72d2f5460e36fda92

Observation f80f7806-56f6-4f3a-adde-ab4af24c94be · outbound

This paper cites Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models.

Potemkin Understanding in Large Language Models Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.380612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.380612Z digest=sha256:1542c67bcc944aadbfb914b7f85b2323f71851483da818efb58e932a93737a54

Observation 0a12bff8-1508-4e41-9956-e1d2c96a5a6b · outbound

This paper cites Adversarial Example Generation with Syntactically Controlled Paraphrase Networks.

Potemkin Understanding in Large Language Models Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.427899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.427899Z digest=sha256:0516834a181aeff55ce411c452f09663cca7ac54499ec6176ffdb845eecd55d2

Observation bd4c4a24-c70e-48cf-a109-fd0779d6db79 · outbound

This paper cites Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models.

Potemkin Understanding in Large Language Models Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:30:33.511891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.469419Z digest=sha256:22889d417776bb059110ea2543afec3115bcfeac7a86cb8604609e047812a7cd

Observation 4ecb5384-98e1-4909-99a9-00cfd77fb317 · outbound

This paper cites S., and Lukasiewicz, T.

Potemkin Understanding in Large Language Models S., and Lukasiewicz, T

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.468299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.557719Z digest=sha256:10523f98978992483a1b0b24913b5674735731af123d60780c3eb3011613a195

Observation 68983c8b-c899-4210-9d69-74e87fd3f5a1 · outbound

This paper cites Consistency Analysis of ChatGPT.

Potemkin Understanding in Large Language Models Consistency Analysis of ChatGPT

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.615032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.615032Z digest=sha256:a60e292672ceaf02ff0ca40f38e540adb4804773b27f095b6ef51abdb9b52b7f

Observation 354c96c6-6b8a-4b97-9c6b-027e9347c2bc · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

Potemkin Understanding in Large Language Models What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.375868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.655859Z digest=sha256:a52463e4c228ffeddde9a71224dc05d23c3902c2135dfb5c193d0dd976ea6fe7

Observation 6c072be9-1d61-4f5e-97e5-5ac8b700aab8 · outbound

This paper cites Dynabench: Rethinking Benchmarking in NLP.

Potemkin Understanding in Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.708684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.708684Z digest=sha256:86018db014af6322427b43925fbd0e20cb9b5e42f89cbf304adab04dc161f158

Observation 0186b315-2370-42f2-bcec-d4c750206225 · outbound

This paper cites a ldchen, S., Binder, A., Montavon, G., Samek, W., and M \.

Potemkin Understanding in Large Language Models a ldchen, S., Binder, A., Montavon, G., Samek, W., and M \

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.273866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:30.785634Z digest=sha256:6f8ef59dfd4e3fb88b1c72596168d8d0223a0043091c65a2317c9d56beae92c1

Observation 61a4b5d1-70c0-452f-accd-f2ac542cbe02 · outbound

This paper cites Benchmarking and Improving Generator-Validator Consistency of Language Models.

Potemkin Understanding in Large Language Models Benchmarking and Improving Generator-Validator Consistency of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.853855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.853855Z digest=sha256:8c93a23ee76c5917f942f5b554412a6244d543964db4c81bdb4f482961d67bc4

Observation 34dd9cdc-8b85-4a39-9792-2ecfad2c4309 · outbound

This paper cites Holistic Evaluation of Language Models.

Potemkin Understanding in Large Language Models Holistic Evaluation of Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.910951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.910951Z digest=sha256:b0b19825c108102ba0fed803e17144c50fd77c3ae739756dbfbf04319907dbd0

Observation 21d94355-3964-424e-89e1-753c74677884 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Potemkin Understanding in Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.965179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.965179Z digest=sha256:7f50235aabb5569d68824b9bb7b0e58b423d4616c10aec94b7d25e5310b97fcf

Observation 2549b930-ffc8-4b3d-934c-a762e4745d65 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

Potemkin Understanding in Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.067924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.067924Z digest=sha256:cc39a535d4467ee2bcd73e71c02b5f590207fe5eb82516ab69293211d529daaf

Observation be411334-3402-45b3-8c6d-ce53abf7f192 · outbound

This paper cites Sycophancy in Large Language Models: Causes and Mitigations.

Potemkin Understanding in Large Language Models Sycophancy in Large Language Models: Causes and Mitigations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.121504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.121504Z digest=sha256:e787f2f55d5a8dd2b8e5bffbcf0f91148939edc1df8f448e1fb695fba2192232

Observation 091c2206-b1ef-4a53-bb76-a814cce90897 · outbound

This paper cites Locating and editing factual associations in gpt.

Potemkin Understanding in Large Language Models Locating and editing factual associations in gpt

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.195017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.176258Z digest=sha256:ba96290825211c348d31db12cad6c355903f32e741477f7212ee8a0df690f18c

Observation 6035467f-3a81-400a-b0d5-cc888ce0ad89 · outbound

This paper cites Fast Model Editing at Scale.

Potemkin Understanding in Large Language Models Fast Model Editing at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.224912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.224912Z digest=sha256:39132ba76ed935fc73f16926e6c3de31f0a8270ff2185ac57026caba2f49d206

Observation f5d35c09-92f8-4221-b27e-3e9898934714 · outbound

This paper cites Why AI is Harder Than We Think.

Potemkin Understanding in Large Language Models Why AI is Harder Than We Think

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.291743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.291743Z digest=sha256:f6e1103b855668ba51f5dca64b883fdf1c67951e4ec35b3c23c7d9fb903fb869

Observation 3f0bdea2-db8d-4924-b014-b67fb3794e51 · outbound

This paper cites D., Bender, E.

Potemkin Understanding in Large Language Models D., Bender, E

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.113972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.379058Z digest=sha256:486f2bda994bd0e5a1de3fcc7698bf24683a37103c9f4bb01a70ecbc5fce8ff8

Observation 4d8bd2d3-03dd-41a4-b40a-9090106f3e7f · outbound

This paper cites Language Models as Knowledge Bases?.

Potemkin Understanding in Large Language Models Language Models as Knowledge Bases?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.430923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.430923Z digest=sha256:0ac72083ba7d94d42b86e311f47c6e7cd58815e41b584d5bbdcdc4ab7155d3ea

Observation c7fccb6c-5ef1-4c84-ad1c-c551ef3259a2 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Potemkin Understanding in Large Language Models Measuring and Narrowing the Compositionality Gap in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.483348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.483348Z digest=sha256:917edf03e0431f69db026a20ecb9e768629294d87de2a25610231af009d8be47

Observation 6c20d737-f59a-4e24-a439-fe41835ade21 · outbound

This paper cites AI and the Everything in the Whole Wide World Benchmark.

Potemkin Understanding in Large Language Models AI and the Everything in the Whole Wide World Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.598191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.598191Z digest=sha256:43319a4ead4547e4e9e84c3c9c78991c058f7d8dd0c20bb0172a7f0322e8e399

Observation 8654d9be-99c3-404f-9f7e-8492fb9d6de5 · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International conference on machine learning.

Potemkin Understanding in Large Language Models Do imagenet classifiers generalize to imagenet? In International conference on machine learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:35.017358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.642476Z digest=sha256:e801cb3865dd0c5d1312e8044139eddf2f381b119ae52d0536b49c3fcd923c83

Observation efe6e90d-51f8-4d07-8a1b-e0d52e412b8e · outbound

This paper cites BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.

Potemkin Understanding in Large Language Models BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.722951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.722951Z digest=sha256:919d944e52669061f3ebb04268156eb306949c300531bf0386166fa6ae4e1d6d

Observation bf318aeb-57fe-4762-ba40-f4bb1662706d · outbound

This paper cites why should i trust you?.

Potemkin Understanding in Large Language Models why should i trust you?

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.941864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.791629Z digest=sha256:b47f77f4a086b30f801b6a1dc7279e12a10b6b1df865d93ae59d2f5aaa6aa75a

Observation 59283c74-da0b-44ab-9b13-74f0e61b7f41 · outbound

This paper cites T., Singh, S., and Guestrin, C.

Potemkin Understanding in Large Language Models T., Singh, S., and Guestrin, C

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.829170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.837348Z digest=sha256:d1d0a68175022a5857d9d9971de4111ca1db35463ee702e189be92560ed97134

Observation d13db075-ea71-430e-a719-e1007ab4d794 · outbound

This paper cites T., Guestrin, C., and Singh, S.

Potemkin Understanding in Large Language Models T., Guestrin, C., and Singh, S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.720084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:31.909318Z digest=sha256:94b0bc8aeba0aa9bebd9a3887bdd491fb4fb545711e99618decf6254a0530928

Observation 1bbb4541-0b04-4f46-83a6-dcf80abbd733 · outbound

This paper cites Beyond Accuracy: Behavioral Testing of NLP models with CheckList.

Potemkin Understanding in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.993054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.993054Z digest=sha256:7d89b168d483f3fc2d860dc82428232ff49b09684512ebd96a1fef90070377a5

Observation 3e4ec196-89d8-4a30-a34d-ea2f8cffbdc8 · outbound

This paper cites Models in the wild: On corruption robustness of neural nlp systems.

Potemkin Understanding in Large Language Models Models in the wild: On corruption robustness of neural nlp systems

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.619733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.060512Z digest=sha256:e8d5032cb37745ed9969e8c81ecd7677dc9343be8271f3bb4b7224d0ececdf4d

Observation 756778bd-8b0b-448d-86bc-517d7c6e4844 · outbound

This paper cites LLMs' Understanding of Natural Language Revealed.

Potemkin Understanding in Large Language Models LLMs' Understanding of Natural Language Revealed

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:30:33.173737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.137478Z digest=sha256:cc31684466e793bc8c96dced753c26dc089b7cc5a2340b944f38ccbd1c38898e

Observation 2cc69fd1-3dd3-41b0-864e-6fd6c07ae10a · outbound

This paper cites everyone wants to do the model work, not the data work.

Potemkin Understanding in Large Language Models everyone wants to do the model work, not the data work

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.533005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.192447Z digest=sha256:d896f46e50684fcab8a4eb848005dfa87c8b2bf3c7a8678e16baa61e69464e4a

Observation 216fad6c-c506-449f-8b23-6b8a3e88fbca · outbound

This paper cites H., Sch \"a rli, N., and Zhou, D.

Potemkin Understanding in Large Language Models H., Sch \"a rli, N., and Zhou, D

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.449212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.275321Z digest=sha256:9d376cfdf1c7ebab9cbd7123d00e520fbd51afa34fc43e527091183cebf4f89b

Observation 016182a3-c9ea-4dc4-b952-cf87cc9e0a17 · outbound

This paper cites and Choi, Y.

Potemkin Understanding in Large Language Models and Choi, Y

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.348695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.351314Z digest=sha256:28e368ee6e26338e2765fb4840aa3a97fd521aba6d071e94de71c632b21c9e61

Observation 14bf4e05-0d71-4d7b-82bf-cfe53b9e85b6 · outbound

This paper cites S., Wei, J., Chung, H.

Potemkin Understanding in Large Language Models S., Wei, J., Chung, H

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.232060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.415105Z digest=sha256:30708c2692c647d0ede27f1181d722c30bd974b8a52b7be3a16e8760b0d72a59

Observation fc60529a-e377-4d7f-bff1-68ebb3aa1c42 · outbound

This paper cites Evaluating the Factual Consistency of Large Language Models Through News Summarization.

Potemkin Understanding in Large Language Models Evaluating the Factual Consistency of Large Language Models Through News Summarization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.462098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.462098Z digest=sha256:121e6c92c3387c644c38708d80eec48e6bedcbfd0d59c7b475142a60fa42f310

Observation e7ad8acb-9e25-4e00-a2f2-9ec75fa9b2f0 · outbound

This paper cites Evaluating the World Model Implicit in a Generative Model.

Potemkin Understanding in Large Language Models Evaluating the World Model Implicit in a Generative Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.551668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.551668Z digest=sha256:e00b7ff2feb9fd7abb86b12e425e9e053d0b7913dcb903e635375b177eff4821

Observation c9c5fe9e-7c31-43b3-91c9-c7e8ec64e601 · outbound

This paper cites Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function.

Potemkin Understanding in Large Language Models Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.594808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.594808Z digest=sha256:b75ff1a6441ff6f4bcccdadda995d04baa255a8c844efc92632377139b08ae72

Observation c157b09a-1b97-49fd-b6bb-facc38aef7bc · outbound

This paper cites On the planning abilities of large language models-a critical investigation.

Potemkin Understanding in Large Language Models On the planning abilities of large language models-a critical investigation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.145704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.648173Z digest=sha256:175ae52e3e34d9dad5242ec63c34a7f6d3de80eaa8abddfa0dcf107e3a177424

Observation ee7bcbd8-a32b-439c-88b7-95a1bbc9be8c · outbound

This paper cites T., Heer, J., and Weld, D.

Potemkin Understanding in Large Language Models T., Heer, J., and Weld, D

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:34.056585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.704793Z digest=sha256:29227d2928daa46d5e1a08fd1424a6b3de2c9c0a431d5cb08ecead4a7581c10e

Observation c4931cb1-6d7c-4e55-a8b2-411fe52458cc · outbound

This paper cites Kformer: Knowledge injection in transformer feed-forward layers.

Potemkin Understanding in Large Language Models Kformer: Knowledge injection in transformer feed-forward layers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:30:33.948182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:30:32.766649Z digest=sha256:ef708c54b4ce49423bb14bef3c9e89f9a6e45b66ed994c655d9730b868b16e34

Observation 360749ae-3ffb-43ba-a9d3-9645dd258d68 · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Potemkin Understanding in Large Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.834986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.834986Z digest=sha256:6de4f4249b81ea473b5fd2d51499dff0865ca3c698efc1715737358faf189f4f

Observation 0468593e-4a93-42a7-8b83-bf4510fa936e · outbound

This paper cites Modifying Memories in Transformer Models.

Potemkin Understanding in Large Language Models Modifying Memories in Transformer Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:32.913920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:32.913920Z digest=sha256:806a3106cde60648628f1561f992f77cf5915eaf9fc35483d30b9cce21127d86

Pith citing papers

Observation e47380ed-7f27-40ba-9c70-a07012ce052f · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Potemkin Understanding in Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.504578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:196596b76280c6d1a6e2af837fd6889e9f53faa8d05d96525a8363c4fb47a672

Observation b920a734-a6e2-4c40-8dfb-4b2310026e8f · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny Potemkin Understanding in Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:14.933677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:14.933677Z digest=sha256:c8356ed76751e89e31be58ca193e5be4e6293d741869aa2ba63d41a3524cfa1b

Observation d3377b6c-ba71-43ca-a1e9-e72554f74392 · inbound

The wall confronting large language models cites this paper.

The wall confronting large language models Potemkin Understanding in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:12:02.086965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:12:02.086965Z digest=sha256:70ec1d64af32978500862acae9bb44883f0591f4324d7b7f0d837b200bf34f0f

Observation 4f151f5c-19ea-4c68-a39f-a7fdbbbd2f34 · inbound

A paradox of AI fluency cites this paper.

A paradox of AI fluency Potemkin Understanding in Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:52.218393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T16:17:48.531790Z digest=sha256:683e379965f09515ae8bd86672f96b63be1af1188dcb94814f729d29a6f82e92

Observation 23d6e3c3-18ed-4570-9836-ee238b71a00a · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Potemkin Understanding in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:d28189e5752bb778e63eb358ecd782e64d3113c4fcd3b28189130ac6161fbb7e