Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2506.02058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02058 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:56:34.350227Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:40:18.133637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:06:01.278099Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved31
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72ab1322-29cf-4754-a082-ab21a584d42e · outbound

This paper cites GPT-4 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.406168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.406168Z digest=sha256:39b717c19880eb752a432f4c67d29ecbaef6d5df05db92b9548a9f18c43f0c9b

Observation 6856b69a-28b4-4b10-85a5-ef242ebea546 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.482658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.482658Z digest=sha256:4b0b339c7359b8f35dbe711f95a9e3f0fdfd8c726efcaf2d4065c35c76e47ad3

Observation 483df34c-5857-4588-a5f8-e33033e45f75 · outbound

This paper cites Claude3.7 Sonnetsystemcard.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Claude3.7 Sonnetsystemcard

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.145924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:29.576094Z digest=sha256:fc59284d601b9385c9d71789ae1ab9cf9f8fd21f5df5785e4a18b0f27f44fad9

Observation 310dca0d-c683-4f36-8c64-c3e4099af996 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.637691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.637691Z digest=sha256:fd6a24c7215647374039c543da458ab1c0bf228d2feb29005cdb1b73d864a421

Observation 688d1427-ca52-42e9-8eeb-76a76e84da90 · outbound

This paper cites Eight things to know about large language models.Critical AI, 2(2), 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Eight things to know about large language models.Critical AI, 2(2), 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.701569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.701569Z digest=sha256:f7c147d847e23b3d783284f558134375303de97209d78931740f616b5452dcd3

Observation 52c15b9c-a07c-4c0d-8f30-eec444771e7d · outbound

This paper cites Language models are few-shot learners.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.768209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.768209Z digest=sha256:fa31972b8db648dd19a0d68789cc2da5dd1da0cf9d39a7211dbbd483c6f9932c

Observation eb8c60f9-0d8a-4f75-b3b3-1ac1710b9d85 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.866099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.866099Z digest=sha256:63cc3b32ebf9525679aad4df3a5b1a2fd615b71eb300cdbe4308cd1450708bab

Observation c7bf8963-f614-4700-8a85-8fd0267fb6ab · outbound

This paper cites Quantifying memorization across neural language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Quantifying memorization across neural language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.124339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:29.927015Z digest=sha256:84da27bfc5cbe8b18a615f12f2eef35913d77b4f0fe4015ca12d400feaf04b5d

Observation acbe91c0-5b91-48fb-a285-504003286703 · outbound

This paper cites How do large language models acquire factual knowledge during pretraining? In Neural Information Processing Systems, 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? How do large language models acquire factual knowledge during pretraining? In Neural Information Processing Systems, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.114386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:30.009946Z digest=sha256:40094ae82b16c5edba906f0d62a2e64950bf5fb0d89ae6eec189e0a11cd14dd1

Observation 3dcb6d7e-dcd2-4478-b92e-73d7c8488979 · outbound

This paper cites A survey on evaluation of large language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A survey on evaluation of large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.089145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.089145Z digest=sha256:f5e255e64a555090676c7f7c483bf981b278ba3911230557cbcdbc77a0e13bc1

Observation c6a89fa4-efc2-4f5e-abc6-b7a8f00ec216 · outbound

This paper cites Estimating the number of species in a stochastic abundance model.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the number of species in a stochastic abundance model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.099450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:30.173634Z digest=sha256:bba0a296185709b377d44de6f7a7fa8cdcfda41eb3152a12d97e3d504713c61e

Observation c10ddbc8-4a61-4e07-a370-281b9cbf1b2d · outbound

This paper cites A new statistical approach for assessing similarity of species composition with incidence and abundance data.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A new statistical approach for assessing similarity of species composition with incidence and abundance data

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.090536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:30.264147Z digest=sha256:c209aa2aec4a264d317701dc84d972488f38ad0dd2676ee3a27c80a82991f840

Observation f6da8e46-f8f9-4d7d-9a97-54c0cf7f2fed · outbound

This paper cites Gonzalez, and Ion Stoica.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Gonzalez, and Ion Stoica

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.081757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:30.373879Z digest=sha256:e0b0451799706fbb65cfe6a5b54e671cf5c970eafc60cc0a4b787833d238dfda

Observation b03eeb4f-30c9-42f9-96c8-d6df7e737740 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.466711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.466711Z digest=sha256:472110e551f16485ab8cbf12dde489e87181abe995ab34462fa2b25bee11665d

Observation 378f9079-eb0f-42d9-8c91-8c2c76031828 · outbound

This paper cites Forgetwhat you know about LLMs evaluations—LLMs are like a chameleon.arXiv preprint arXiv:2502.07445, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Forgetwhat you know about LLMs evaluations—LLMs are like a chameleon.arXiv preprint arXiv:2502.07445, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.563192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.563192Z digest=sha256:2ec8466bc07276363b06b20d0df8f8e9ead2db5c38489a4b8c5984a28b2157de

Observation 48ea9d4c-b2af-4788-8e7a-be70489ca6f4 · outbound

This paper cites Springer US, Boston, MA, 2009.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Springer US, Boston, MA, 2009

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.691748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.691748Z digest=sha256:3a7eeb5a9e08744aee84afd3c6dbf6c8a045338d75a5f68ca7039d5b03c6f252

Observation 1127a67f-eacf-47b3-8e4f-5f31f4a9b94c · outbound

This paper cites CURIE: Evaluating LLMs on multitask scientific long-context understanding and reasoning.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? CURIE: Evaluating LLMs on multitask scientific long-context understanding and reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.072802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:30.800214Z digest=sha256:b706319163b5f304bf16ef37eb488e3602013c5274995cdb969ec271c1b18153

Observation 3dd36863-3d46-41a1-b651-f92f6d11ca24 · outbound

This paper cites Data science at the singularity.Harvard Data Science Review, 6(1), 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Data science at the singularity.Harvard Data Science Review, 6(1), 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.902979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.902979Z digest=sha256:dd59c15405fa7b97cbcb68f42d9475b36d6dbf0cd5480f8249b19bf4ebc95a88

Observation b88b454a-e470-4841-8662-022324c9d8d5 · outbound

This paper cites Estimating the number of unseen species: How many words did Shakespeare know?Biometrika, 63(3):435–447, 1976.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the number of unseen species: How many words did Shakespeare know?Biometrika, 63(3):435–447, 1976

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.057578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:30.995955Z digest=sha256:b5727794d467725f5bbd60b95ac98cf71d80f22542316fe52b1402ee71b30765

Observation fec3978d-2ff7-424c-a75c-b4eb4fbd38ac · outbound

This paper cites Near-optimal estimation of the unseen under regularly varying tail populations.Bernoulli, 29(4):3423–3442, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Near-optimal estimation of the unseen under regularly varying tail populations.Bernoulli, 29(4):3423–3442, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.047540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:31.108281Z digest=sha256:591ba4f3054bee206bafb0e8e410e8bc24fd6ae0c05e1833a6e24460f739dc80

Observation ca5772c1-a696-4352-a239-2cbaa2b9bfed · outbound

This paper cites Good-turing frequency estimation without tears.Journal of quantitative linguistics, 2(3):217–237, 1995.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Good-turing frequency estimation without tears.Journal of quantitative linguistics, 2(3):217–237, 1995

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.201554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.201554Z digest=sha256:19e645fc2cf0cea3816aa48c3083301e5e0afeb410562c343a29c90555ebd3a5

Observation 2f94198c-03f9-46b7-83f1-1ac2be5258f3 · outbound

This paper cites The population frequencies of species and the estimation of population parameters.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The population frequencies of species and the estimation of population parameters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.031401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:31.316125Z digest=sha256:741b617f0a69b4a018e9f97923786a5f09f088efb0228f63f06bd21d3f4598f1

Observation f11c90b2-fe19-45a4-9762-1131b0048e9b · outbound

This paper cites Estimating knowledge in large language models without generating a single token.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating knowledge in large language models without generating a single token

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.021476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:31.399782Z digest=sha256:2564a55d5d79e2d2a97c8e741e10d02f8b7797cab879fc14499f7ff68bcb5a47

Observation f17a9559-781b-4bc9-8727-ba3f6e3ed67c · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.486703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.486703Z digest=sha256:ab240036e3ee85390583960e156e806ebbf9bdc1dd0fe63cc04fc07b7ad3f838

Observation aae552b3-fbaa-4cbf-9c85-4b0dfad527ba · outbound

This paper cites Optimal prediction of the number of unseen species with multiplicity.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Optimal prediction of the number of unseen species with multiplicity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.011026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:31.589683Z digest=sha256:c6980472d2d5279879c3cc3f5458a06f902383c76815577a9ed983a7532f5a41

Observation e1b4f6e4-a378-4ed5-83a4-01e165bbcbb1 · outbound

This paper cites Measuring massive multitask language understanding.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Measuring massive multitask language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.701733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.701733Z digest=sha256:7340ba6e2f97ae3275d1558f8c702882d12974f72268dcb5889f3d5b94f46e22

Observation 2d78d870-0ef7-45a6-8018-4b91dc929097 · outbound

This paper cites The curious case of neural text degeneration.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The curious case of neural text degeneration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.806850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.806850Z digest=sha256:faf08f72bcb0472dfa3c3e4e037be2944df1b60118e7f7987014ce556aa155f4

Observation 3369593b-ec38-4dc2-93f2-35b1cbc63d97 · outbound

This paper cites GPT-4o System Card.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? GPT-4o System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.905024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.905024Z digest=sha256:7ca963adcae4ba5ec3d2a6b2470b5351b9bd1951ec7abd87491731f2bc0a4c75

Observation 8b7ece32-5645-4610-a2fb-3ebc3a7c75dc · outbound

This paper cites Mistral 7B.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.057825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.057825Z digest=sha256:80433c46fa9134bded48995db1751f33cdcea106a3081733a49195d90e7cd216

Observation e6655ea8-530c-4da3-ae41-b5589d52714f · outbound

This paper cites Calibrated language models must hallucinate.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Calibrated language models must hallucinate

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.989799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.156600Z digest=sha256:cfb75f84449150b3adab2838c6f690f805e3291f9ad7750656ec1393e8c20bf3

Observation b677d7cb-2225-45af-b630-69fe5edaf29d · outbound

This paper cites Line of duty: Evaluating LLM self-knowledge via consistency in feasibility boundaries.arXiv preprint arXiv:2503.11256, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Line of duty: Evaluating LLM self-knowledge via consistency in feasibility boundaries.arXiv preprint arXiv:2503.11256, 2025

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-07T11:56:34.654800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.274696Z digest=sha256:378fc0fbe6d0e224f89df6fe5837387952c93dfe6265c6b2694704f0b87010d1

Observation aec70103-f245-46c7-a6cd-500493c24cc3 · outbound

This paper cites Too many AIs.https://dev.to/leeaao/too-many-ais-24nb, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Too many AIs.https://dev.to/leeaao/too-many-ais-24nb, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.979147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.398334Z digest=sha256:d725f087bd8a8565cd4abb258dc8bcf56daa840460afe4f98d45e11cf48f4050

Observation 57c732cf-9d26-4733-8cec-29c15d6a59d4 · outbound

This paper cites BioASQ-QA: A manually curated corpus for biomedical question answering.Scientific Data, 10 (1):170, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? BioASQ-QA: A manually curated corpus for biomedical question answering.Scientific Data, 10 (1):170, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.969869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.490401Z digest=sha256:18aad2188f5d117e62cb5020e55cbb460a1c8857b90cf0e8f13130be2882f0df

Observation 1fa5d269-dc34-47e8-bc34-dcdc8e918390 · outbound

This paper cites How pre-trained language models capture factual knowledge? A causal-inspired analysis.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? How pre-trained language models capture factual knowledge? A causal-inspired analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.960897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.562356Z digest=sha256:33678fd917c41cd115e0ba1b8b7444228dfd10363bbbb97d3b9e26b8fcaf1e4e

Observation 95f36e0b-4d3d-4373-b917-47032910befc · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? ROUGE: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.951469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.665033Z digest=sha256:5620513482defe6c739dc83ed2ed2771e46edb6e247bb5938552e5176ff5942e

Observation a9b5a12a-3641-4fcb-9030-656528f8c9ac · outbound

This paper cites DeepSeek-V3 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? DeepSeek-V3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.742702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.742702Z digest=sha256:cbe816dec82d8d3aed32549049c842fd830a6556c9bface2c11fe27fe3306f66

Observation 0ab7b408-c554-4834-8f88-2094453ab49a · outbound

This paper cites Feder Cooper, Daphne Ippolito, Christopher A.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Feder Cooper, Daphne Ippolito, Christopher A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.942557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.831463Z digest=sha256:37207a878e4eda5e12b8039133e2db4d7567ffc118a8c95a7256b87976e1aaa5

Observation 9bdadb55-1361-4762-aa38-54a9fcac97a8 · outbound

This paper cites ChatGPT-3.5-turbo.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? ChatGPT-3.5-turbo

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.933728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:32.939632Z digest=sha256:fbe9c8ee0cd48ab3572e6c2cb5a2eb80c406c3d269fccded6d519e7c395a1dd8

Observation 85ce83e6-eff3-4976-9842-f6f8c64f84f4 · outbound

This paper cites Competitive distribution estimation: Why is Good-Turing good.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Competitive distribution estimation: Why is Good-Turing good

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.925398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.057660Z digest=sha256:d39b210b7869ba6dfbb1bdf60ac1a9905f0f9c082ab158db7ff1f0e69d2bc987

Observation f17f5189-d0c5-46f6-8727-8bcb583d3a86 · outbound

This paper cites Always Good Turing: Asymptotically optimal probability estimation.Science, 302(5644):427–431, 2003.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Always Good Turing: Asymptotically optimal probability estimation.Science, 302(5644):427–431, 2003

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.916852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.170593Z digest=sha256:4e607bbb8005c38c7df89c9e78cedbb9848294b43711a18d600a703235f55b7b

Observation f2c6e254-e81a-4a3b-8fbc-06a456d11fd1 · outbound

This paper cites Optimal prediction of the number of unseen species.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Optimal prediction of the number of unseen species

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.907642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.261347Z digest=sha256:f18af2bdb8f3232ddfc2497e1049427533220c73e1cf5e2cadff7547db0b7329

Observation ce2a9fe4-7ca7-4a92-98d4-b5a1a592c3c0 · outbound

This paper cites BLEU: A method for automatic evaluation of machine translation.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? BLEU: A method for automatic evaluation of machine translation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.897789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.392631Z digest=sha256:2a79f0b432bf5f6080f424ad9635b6edbeefed0390efa272e00e379825bd90dd

Observation 0faef811-19d8-4ae0-9219-45c3b109bcb2 · outbound

This paper cites an unresolved cited work.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:56:34.888538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.542099Z digest=sha256:c9eda5f997d34fe1be0964c97888085beb3d617a36acf667f4688ed5ab26f024

Observation 68e09110-55f4-421b-9086-8ebd96634a3c · outbound

This paper cites Humanity's Last Exam.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Humanity's Last Exam

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.693512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.693512Z digest=sha256:4f773cdb3bd11cc1b4ae8837ff638d3ab4307dbf1e96903615f84e600aafa558

Observation 900d1761-f576-45ef-b239-387891575e95 · outbound

This paper cites Do large language models know how much they know?arXiv preprint arXiv:2502.19573, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Do large language models know how much they know?arXiv preprint arXiv:2502.19573, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.840301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.840301Z digest=sha256:3f07678f4562a9f1fd48610410d51606fcd820aec9d3cb1ea317f75fdf5aba80

Observation 96fb3030-8cd4-4bf9-90ff-ba9b27161418 · outbound

This paper cites AI and the everything in the whole wide world benchmark.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? AI and the everything in the whole wide world benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.879529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.927131Z digest=sha256:75c7d8cbb65b423c80fbd2a96041c1c2595f67156f38af725d472d1b8b277969

Observation 4b1c0bfa-8e40-4db1-82b9-0aa6c63ca32c · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.870380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:33.972092Z digest=sha256:dd42f26ebb8185ed93a8963c066f19b9f7973caf4f0b5d2179de456d8702b774

Observation af1bb3c2-8c34-48f4-ba7e-f3450f84ede4 · outbound

This paper cites NeurIPS 2023 LLM Efficiency Fine-tuning Competition.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? NeurIPS 2023 LLM Efficiency Fine-tuning Competition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.073212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.073212Z digest=sha256:4cbec0aca7bf63307fae3858557477bb2f51c4711ec50baeda6fe534b6b9aef9

Observation fadec12e-09ec-48bd-bdad-4678fb372d40 · outbound

This paper cites Human disease ontology 2022 update.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Human disease ontology 2022 update

Reference 49

Resolution
verified exact
doi, observed 2026-08-07T11:56:34.393035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.173184Z digest=sha256:b87b607779b41766185894848b091d2150dcf44038c18f60d870e16c8737c155

Observation bfa37a32-4556-4365-bde5-612215903355 · outbound

This paper cites Auto- Prompt: Eliciting knowledge from language models with automatically generated prompts.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Auto- Prompt: Eliciting knowledge from language models with automatically generated prompts

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.860670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.298457Z digest=sha256:cb52bbb1ffc6d2a717d11e7461b816d9bc97ae60934f307ee7ab34ced2478ae4

Observation 4278205c-6193-4b12-809a-f2a1b2779fd5 · outbound

This paper cites Welcome to the era of experience.Google AI, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Welcome to the era of experience.Google AI, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.304140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.304140Z digest=sha256:3e8453a247268c086849e71e90e129331e9cf0337062a4a4d34e5a687088d7d7

Observation cb6f8d71-aa20-4602-a8c0-b521e391c82c · outbound

This paper cites The Leaderboard Illusion.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The Leaderboard Illusion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.307356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.307356Z digest=sha256:4de3f78175662bb5fb6114fe2e2b741ba5f80c78abd526140a4b55f37e461654

Observation 0714b062-436d-42fa-8830-9680c3342321 · outbound

This paper cites Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.311072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.311072Z digest=sha256:89bba5b60f4c572532644704524dda4e087b2bf0167b7f23065c173e3f13d16d

Observation 63154cd9-7896-4b3b-bc9f-231e38059ce7 · outbound

This paper cites The bitter lesson.Incomplete Ideas (blog), 13(1):38, 2019.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The bitter lesson.Incomplete Ideas (blog), 13(1):38, 2019

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.314529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.314529Z digest=sha256:dc44ade7f4c33578119338744dbca62e00c683c41d02e1f1a36fd91a0dfffdbd

Observation cdf82a61-b4c1-4319-99ba-e27ea41fb45c · outbound

This paper cites Galactica: A Large Language Model for Science.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Galactica: A Large Language Model for Science

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.317573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.317573Z digest=sha256:04fc35a9d29f4769135fedb9afb5b05ccf611c5f1dd8660571c01e18067b36e3

Observation bf256b5f-6e75-449f-bc4e-52fb3deaba4d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.320882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.320882Z digest=sha256:58e26aace53cebb227868612ea99872c6e4c353154571d7cba22b3e9433f53fa

Observation 18c493fd-f9b1-4e8b-9689-78de7603398e · outbound

This paper cites an unresolved cited work.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Unresolved cited work

Reference 57

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T11:56:34.382001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.324022Z digest=sha256:a860282223f22a28eef2c8673a4ef8120c8e3487c4738d7835bfc75553f66ed5

Observation 64e7e780-3939-43a2-953a-1324e6ab1c40 · outbound

This paper cites Language Models are Open Knowledge Graphs.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Language Models are Open Knowledge Graphs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.327247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.327247Z digest=sha256:7d9a3d28c87f8570858d42e395e58bb13e1d5b044ed869a82e5ca68ec9590542

Observation 60d4733b-45b1-4bef-880f-000b5a15117c · outbound

This paper cites Can AI Be as Creative as Humans?.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Can AI Be as Creative as Humans?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.330717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.330717Z digest=sha256:11e2909f5161e675fe5c8aa45e688f8311904cd477fb39b24c27bbb34f99bc1c

Observation 5d323f19-07f3-41a9-9f95-676aaa7aae99 · outbound

This paper cites Chi, Quoc V.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Chi, Quoc V

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.830980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.334226Z digest=sha256:3f30f8562fcb1924664d97db3b8006d73648329275f87fef8bd523d8db14e02f

Observation 40608028-8344-4c9d-bbfe-42717f898ea3 · outbound

This paper cites Estimating the probabilities of rare outputs in language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the probabilities of rare outputs in language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.822545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.337670Z digest=sha256:893a1e4b698529d413048b41f3087177ea5fc5cd3b6bd33bcd9e4fb90f4a213e

Observation 8f4f37ea-0c4b-4110-a534-f91081bc8a5e · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Qwen2.5 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.340659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.340659Z digest=sha256:35d0df083b3b9d49a0244cb3e0d30bb48d21fe2d7170ee14e0988f0f29a6f460

Observation 893d72c8-1572-424b-9370-94f79a7c44b2 · outbound

This paper cites A careful examination of large language model performance on grade school arithmetic.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A careful examination of large language model performance on grade school arithmetic

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.812963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.343378Z digest=sha256:d8639bfb46102933205c13a4443b9375ec65e9555ed8fcb23c40982bc18cf380

Observation 88fcaddc-d668-4a60-87fd-a06a821f89a5 · outbound

This paper cites Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:56:34.413661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.346706Z digest=sha256:ace36b7bf90cb9adc229be2393eee1a9384e7fb164e7e2878228ec1953df2798

Observation 29d200ca-a50b-460f-8047-70ddb01d209f · outbound

This paper cites Test your knowledge of mathematical theorems by listing 20 theorem names, separated by commas.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Test your knowledge of mathematical theorems by listing 20 theorem names, separated by commas

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:56:34.803679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:56:34.350227Z digest=sha256:1c1d99baa12d6e39d403af4667974517d7873cabf22c2cf39e68bc21efb0994c

Pith citing papers

Observation 4f39711d-05df-4367-bdf0-2987ca8ae6db · inbound

UCS: Estimating Unseen Coverage for Improved In-Context Learning cites this paper.

UCS: Estimating Unseen Coverage for Improved In-Context Learning Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:01.283344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:40:18.133637Z digest=sha256:33a5dc9ebdafdac767ef09636d1565692d9992928370773a4a449871b78dc3b3