Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2506.02058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02058 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:56:34.350227Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:40:18.133637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:06:01.278099Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved31
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72ab1322-29cf-4754-a082-ab21a584d42e · outbound

This paper cites GPT-4 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.406168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.406168Z digest=sha256:5e1e26dcdc0ff5acd8c563a42f603086e7d0382c81114d0fdae3a77839f0e70c

Observation 6856b69a-28b4-4b10-85a5-ef242ebea546 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.482658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.482658Z digest=sha256:feed9db94f0ebadb7b1b34428dfa22bcaa9c822d4713478a50e8205f80cc9114

Observation 483df34c-5857-4588-a5f8-e33033e45f75 · outbound

This paper cites Claude3.7 Sonnetsystemcard.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Claude3.7 Sonnetsystemcard

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.145924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:29.576094Z digest=sha256:eeef3067276074de245be12f6eb916c3c7278690a31f81a134256e026b4e9f6e

Observation 310dca0d-c683-4f36-8c64-c3e4099af996 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.637691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.637691Z digest=sha256:9da0ef3a35498ca5c4f74b7c2773f428513cb6de653ad4f5d97573ccbfc69ecc

Observation 688d1427-ca52-42e9-8eeb-76a76e84da90 · outbound

This paper cites Eight things to know about large language models.Critical AI, 2(2), 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Eight things to know about large language models.Critical AI, 2(2), 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.701569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.701569Z digest=sha256:6e1c5864d5657666ce92ffe0f9c6d74dce4bf99ffdc1874823d126da4646012f

Observation 52c15b9c-a07c-4c0d-8f30-eec444771e7d · outbound

This paper cites Language models are few-shot learners.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.768209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.768209Z digest=sha256:9635c3290d56488d5791c5a69d29ee65aeb79fb30078a8ffdaa5e8e6837f38cc

Observation eb8c60f9-0d8a-4f75-b3b3-1ac1710b9d85 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.866099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.866099Z digest=sha256:74992a15a8ad4fcc05b4ff8e58c6602edfa8db01089149a77888016594f04557

Observation c7bf8963-f614-4700-8a85-8fd0267fb6ab · outbound

This paper cites Quantifying memorization across neural language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Quantifying memorization across neural language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.124339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:29.927015Z digest=sha256:b3a29bed2861b7eba0e7cac97c7817e7e7229d915b2519710902371b694e83f3

Observation acbe91c0-5b91-48fb-a285-504003286703 · outbound

This paper cites How do large language models acquire factual knowledge during pretraining? In Neural Information Processing Systems, 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? How do large language models acquire factual knowledge during pretraining? In Neural Information Processing Systems, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.114386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:30.009946Z digest=sha256:5fd5d2ce782c42859e5f31341e2c722ae414167c821e39016da0d3e1a028ec85

Observation 3dcb6d7e-dcd2-4478-b92e-73d7c8488979 · outbound

This paper cites A survey on evaluation of large language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A survey on evaluation of large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.089145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.089145Z digest=sha256:0860ba65aa085bd4957101b5d902880bd53e7c3cbdab15eea6f9b1127c701d12

Observation c6a89fa4-efc2-4f5e-abc6-b7a8f00ec216 · outbound

This paper cites Estimating the number of species in a stochastic abundance model.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the number of species in a stochastic abundance model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.099450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:30.173634Z digest=sha256:d3d72646f5cb0d4edd65310cdbd6f8f36dccce500f40adbc3a81cdf112e660d2

Observation c10ddbc8-4a61-4e07-a370-281b9cbf1b2d · outbound

This paper cites A new statistical approach for assessing similarity of species composition with incidence and abundance data.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A new statistical approach for assessing similarity of species composition with incidence and abundance data

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.090536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:30.264147Z digest=sha256:1f25da793a86f584d8bb0d51e4f370ee436e1d80966f0db38276ee519a034b04

Observation f6da8e46-f8f9-4d7d-9a97-54c0cf7f2fed · outbound

This paper cites Gonzalez, and Ion Stoica.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Gonzalez, and Ion Stoica

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.081757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:30.373879Z digest=sha256:dbd04e4863bc845f395eec655892ee43edaa459a02be8c2315ce3708f204e06f

Observation b03eeb4f-30c9-42f9-96c8-d6df7e737740 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.466711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.466711Z digest=sha256:f3d82d914214a56f5fa9290e76b62f09b42cd46459f1d08307444d8fb23ff45f

Observation 378f9079-eb0f-42d9-8c91-8c2c76031828 · outbound

This paper cites Forgetwhat you know about LLMs evaluations—LLMs are like a chameleon.arXiv preprint arXiv:2502.07445, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Forgetwhat you know about LLMs evaluations—LLMs are like a chameleon.arXiv preprint arXiv:2502.07445, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.563192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.563192Z digest=sha256:e395036b512ae6833b5654580249c64ad0dd778404a1e2038332a79ea7bde5f9

Observation 48ea9d4c-b2af-4788-8e7a-be70489ca6f4 · outbound

This paper cites Springer US, Boston, MA, 2009.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Springer US, Boston, MA, 2009

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.691748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.691748Z digest=sha256:342c8a79b1e38e55acb12d8578a45b368cfd7e1b2643e1db75ad20dd86dde850

Observation 1127a67f-eacf-47b3-8e4f-5f31f4a9b94c · outbound

This paper cites CURIE: Evaluating LLMs on multitask scientific long-context understanding and reasoning.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? CURIE: Evaluating LLMs on multitask scientific long-context understanding and reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.072802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:30.800214Z digest=sha256:f0df134fa1daf0b2bae778959e42d562d76502a3d6931b27ba52180cd5d25e2b

Observation 3dd36863-3d46-41a1-b651-f92f6d11ca24 · outbound

This paper cites Data science at the singularity.Harvard Data Science Review, 6(1), 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Data science at the singularity.Harvard Data Science Review, 6(1), 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.902979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.902979Z digest=sha256:d8dd7ab8975cd05dfe2868b48a84020446e20cc81388e95e6060807077c142ad

Observation b88b454a-e470-4841-8662-022324c9d8d5 · outbound

This paper cites Estimating the number of unseen species: How many words did Shakespeare know?Biometrika, 63(3):435–447, 1976.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the number of unseen species: How many words did Shakespeare know?Biometrika, 63(3):435–447, 1976

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.057578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:30.995955Z digest=sha256:86ce53bf85fff34553df3a0ddc8dd70b5edee287c2d9f1903712876150795a33

Observation fec3978d-2ff7-424c-a75c-b4eb4fbd38ac · outbound

This paper cites Near-optimal estimation of the unseen under regularly varying tail populations.Bernoulli, 29(4):3423–3442, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Near-optimal estimation of the unseen under regularly varying tail populations.Bernoulli, 29(4):3423–3442, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.047540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:31.108281Z digest=sha256:216ee59fa4d5a68cab9dd65cf22a334846935e117fcd3a00d6030c3aeca1d37c

Observation ca5772c1-a696-4352-a239-2cbaa2b9bfed · outbound

This paper cites Good-turing frequency estimation without tears.Journal of quantitative linguistics, 2(3):217–237, 1995.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Good-turing frequency estimation without tears.Journal of quantitative linguistics, 2(3):217–237, 1995

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.201554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.201554Z digest=sha256:f23d92885cb33891750ed7193713d739e46be03e7957bce3a57884dbc7cf8fdb

Observation 2f94198c-03f9-46b7-83f1-1ac2be5258f3 · outbound

This paper cites The population frequencies of species and the estimation of population parameters.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The population frequencies of species and the estimation of population parameters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.031401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:31.316125Z digest=sha256:bdef126e92c89a9c9de2087b44b563765c35e0c39dadd78f759840f4d853ab0c

Observation f11c90b2-fe19-45a4-9762-1131b0048e9b · outbound

This paper cites Estimating knowledge in large language models without generating a single token.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating knowledge in large language models without generating a single token

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.021476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:31.399782Z digest=sha256:d332d1a9b5dba8765b51e05582bb4f5b9ad8b8607f288575a09c0341ac03759c

Observation f17a9559-781b-4bc9-8727-ba3f6e3ed67c · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.486703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.486703Z digest=sha256:770ad358e14da09e5dfea8e5bda686040ea6ea87ffb48f3ac0f9de7738bfbb3d

Observation aae552b3-fbaa-4cbf-9c85-4b0dfad527ba · outbound

This paper cites Optimal prediction of the number of unseen species with multiplicity.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Optimal prediction of the number of unseen species with multiplicity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.011026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:31.589683Z digest=sha256:afd3f1f60d0f1caddfd97a6d1a88db8b4684b1b7f76a2f82811e21797e8714f5

Observation e1b4f6e4-a378-4ed5-83a4-01e165bbcbb1 · outbound

This paper cites Measuring massive multitask language understanding.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Measuring massive multitask language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.701733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.701733Z digest=sha256:50ee484be592bd326beb454a155fbac77835ce3a4a9dcf8202224787de2e8770

Observation 2d78d870-0ef7-45a6-8018-4b91dc929097 · outbound

This paper cites The curious case of neural text degeneration.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The curious case of neural text degeneration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.806850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.806850Z digest=sha256:3c0a6bdecd33ab5a05cbbc063276ed6882643d5c9b1a6b5efc1d06d05d2872ec

Observation 3369593b-ec38-4dc2-93f2-35b1cbc63d97 · outbound

This paper cites GPT-4o System Card.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? GPT-4o System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.905024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.905024Z digest=sha256:eb72e6a61567ed7bafb3de04bc2ac779a8ec6e6f025605e71d450bfeeaba4494

Observation 8b7ece32-5645-4610-a2fb-3ebc3a7c75dc · outbound

This paper cites Mistral 7B.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.057825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.057825Z digest=sha256:fbbe3f116bf4c13636e9e3fe0bf450daf00096e1338191143a12aedf3481b473

Observation e6655ea8-530c-4da3-ae41-b5589d52714f · outbound

This paper cites Calibrated language models must hallucinate.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Calibrated language models must hallucinate

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.989799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.156600Z digest=sha256:46f37fd59486c82801d19d343827a427d497820334ed9768f835cabb75a9107d

Observation b677d7cb-2225-45af-b630-69fe5edaf29d · outbound

This paper cites Line of duty: Evaluating LLM self-knowledge via consistency in feasibility boundaries.arXiv preprint arXiv:2503.11256, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Line of duty: Evaluating LLM self-knowledge via consistency in feasibility boundaries.arXiv preprint arXiv:2503.11256, 2025

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-07T11:56:34.654800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.274696Z digest=sha256:4e337120d5f5fc9c236303805a112172e9f356a43e2d8660ef1e8419e2eed313

Observation aec70103-f245-46c7-a6cd-500493c24cc3 · outbound

This paper cites Too many AIs.https://dev.to/leeaao/too-many-ais-24nb, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Too many AIs.https://dev.to/leeaao/too-many-ais-24nb, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.979147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.398334Z digest=sha256:e99aec70afa027803d5384a57252e4a4bf4dac68a5bf9a4fbd358d84e4d83f63

Observation 57c732cf-9d26-4733-8cec-29c15d6a59d4 · outbound

This paper cites BioASQ-QA: A manually curated corpus for biomedical question answering.Scientific Data, 10 (1):170, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? BioASQ-QA: A manually curated corpus for biomedical question answering.Scientific Data, 10 (1):170, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.969869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.490401Z digest=sha256:06cc6740830e6a01eca8f6e28d68f323921305a776ce522f8715d7a8eac18438

Observation 1fa5d269-dc34-47e8-bc34-dcdc8e918390 · outbound

This paper cites How pre-trained language models capture factual knowledge? A causal-inspired analysis.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? How pre-trained language models capture factual knowledge? A causal-inspired analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.960897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.562356Z digest=sha256:face9770f145db95178fa07eab4a14a1b37615395144ba06b7cd709238a1c9d9

Observation 95f36e0b-4d3d-4373-b917-47032910befc · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? ROUGE: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.951469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.665033Z digest=sha256:07f3b7e68fa71aacd161a35d24ff8bb1bcd0479d41961758cad1f2d6bfbb77c5

Observation a9b5a12a-3641-4fcb-9030-656528f8c9ac · outbound

This paper cites DeepSeek-V3 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? DeepSeek-V3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.742702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.742702Z digest=sha256:36c14c494db93aa732cdf719ec95ddf228696d0b18faed2ede3d6d79e072400e

Observation 0ab7b408-c554-4834-8f88-2094453ab49a · outbound

This paper cites Feder Cooper, Daphne Ippolito, Christopher A.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Feder Cooper, Daphne Ippolito, Christopher A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.942557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.831463Z digest=sha256:ed8271e9d94a362373b289ab2bd4047de4ed8516f5d148320e3eb5a268d7e186

Observation 9bdadb55-1361-4762-aa38-54a9fcac97a8 · outbound

This paper cites ChatGPT-3.5-turbo.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? ChatGPT-3.5-turbo

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.933728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:32.939632Z digest=sha256:c2ea078889484afed8b63b6522e54ba1a0b2de9b1e37fad6ee3d6105fe95a13f

Observation 85ce83e6-eff3-4976-9842-f6f8c64f84f4 · outbound

This paper cites Competitive distribution estimation: Why is Good-Turing good.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Competitive distribution estimation: Why is Good-Turing good

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.925398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.057660Z digest=sha256:2c8129085a8726f6e3f531013e813b65b031d0420c39204fb752da562d87357f

Observation f17f5189-d0c5-46f6-8727-8bcb583d3a86 · outbound

This paper cites Always Good Turing: Asymptotically optimal probability estimation.Science, 302(5644):427–431, 2003.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Always Good Turing: Asymptotically optimal probability estimation.Science, 302(5644):427–431, 2003

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.916852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.170593Z digest=sha256:8fed2a9f0a8b81e842620dc4c2e8a91614240adb464be1c19e80535a909c0cee

Observation f2c6e254-e81a-4a3b-8fbc-06a456d11fd1 · outbound

This paper cites Optimal prediction of the number of unseen species.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Optimal prediction of the number of unseen species

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.907642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.261347Z digest=sha256:5f5700360930305da68d0ff55c003afd92b65d34479fe749621d66f003fb6656

Observation ce2a9fe4-7ca7-4a92-98d4-b5a1a592c3c0 · outbound

This paper cites BLEU: A method for automatic evaluation of machine translation.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? BLEU: A method for automatic evaluation of machine translation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.897789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.392631Z digest=sha256:8005dd64b840bc0f3b5cf475e2d3707872a99cef48756dca9bdf76693044292c

Observation 0faef811-19d8-4ae0-9219-45c3b109bcb2 · outbound

This paper cites an unresolved cited work.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:56:34.888538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.542099Z digest=sha256:bd1429222db36e2b7cd0248dbfb6ded854fc392a89cae2446d35b96a3fecf838

Observation 68e09110-55f4-421b-9086-8ebd96634a3c · outbound

This paper cites Humanity's Last Exam.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Humanity's Last Exam

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.693512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.693512Z digest=sha256:dbeedae9a196fc5eb34581ed1d7e32e7b1014550aa51747b9ce703f9c9e88ac0

Observation 900d1761-f576-45ef-b239-387891575e95 · outbound

This paper cites Do large language models know how much they know?arXiv preprint arXiv:2502.19573, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Do large language models know how much they know?arXiv preprint arXiv:2502.19573, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.840301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.840301Z digest=sha256:96f1e797c8469596a5cb2fc9163f6c95fb9afde184f12c4526126e1c5160e734

Observation 96fb3030-8cd4-4bf9-90ff-ba9b27161418 · outbound

This paper cites AI and the everything in the whole wide world benchmark.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? AI and the everything in the whole wide world benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.879529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.927131Z digest=sha256:ce0ded2c86a5ac7cb887ebeac69de3b26e7ccea96326bed0738bf1ca7bd8a479

Observation 4b1c0bfa-8e40-4db1-82b9-0aa6c63ca32c · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.870380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:33.972092Z digest=sha256:2d15ec34fa138bb38d1d6812aa455987e98bba3eb6860cb2a0058ef2e980a66b

Observation af1bb3c2-8c34-48f4-ba7e-f3450f84ede4 · outbound

This paper cites NeurIPS 2023 LLM Efficiency Fine-tuning Competition.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? NeurIPS 2023 LLM Efficiency Fine-tuning Competition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.073212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.073212Z digest=sha256:00d2b1a74dc2ea72a070f150f2247c72ac789291fe66b66412a69b3b0fe96ab3

Observation fadec12e-09ec-48bd-bdad-4678fb372d40 · outbound

This paper cites Human disease ontology 2022 update.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Human disease ontology 2022 update

Reference 49

Resolution
verified exact
doi, observed 2026-08-07T11:56:34.393035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.173184Z digest=sha256:af8d8144ba692665ed0595067e05624d6c4edfb83208f6d825b23878e099cd8d

Observation bfa37a32-4556-4365-bde5-612215903355 · outbound

This paper cites Auto- Prompt: Eliciting knowledge from language models with automatically generated prompts.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Auto- Prompt: Eliciting knowledge from language models with automatically generated prompts

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.860670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.298457Z digest=sha256:5ada286526750ca734068389f703aa5ebe763cc33cd6c644d8d6595fe80fce30

Observation 4278205c-6193-4b12-809a-f2a1b2779fd5 · outbound

This paper cites Welcome to the era of experience.Google AI, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Welcome to the era of experience.Google AI, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.304140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.304140Z digest=sha256:43999fbc34426f51eecbf099bd2e5b718f4beb22de307e0afbe1a877816f1b6c

Observation cb6f8d71-aa20-4602-a8c0-b521e391c82c · outbound

This paper cites The Leaderboard Illusion.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The Leaderboard Illusion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.307356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.307356Z digest=sha256:a957f66c178263c60bd417909b57e893b5f2a907f614391814cee501c0cd84e7

Observation 0714b062-436d-42fa-8830-9680c3342321 · outbound

This paper cites Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.311072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.311072Z digest=sha256:478b1aa6f49215b65506d820e3bf030880a13bb6efa4931542c4d2aec62525f3

Observation 63154cd9-7896-4b3b-bc9f-231e38059ce7 · outbound

This paper cites The bitter lesson.Incomplete Ideas (blog), 13(1):38, 2019.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The bitter lesson.Incomplete Ideas (blog), 13(1):38, 2019

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.314529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.314529Z digest=sha256:d59094c9805f5f7afe79c2395995f8393ff910e4787323508fe0ffca3e938279

Observation cdf82a61-b4c1-4319-99ba-e27ea41fb45c · outbound

This paper cites Galactica: A Large Language Model for Science.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Galactica: A Large Language Model for Science

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.317573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.317573Z digest=sha256:5240b0630670c0cf1d295b90c3c96f8ac669a2e0f8b365d726f9252ea88596e0

Observation bf256b5f-6e75-449f-bc4e-52fb3deaba4d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.320882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.320882Z digest=sha256:71c27bc57376fe669340ba814bd0d92b5ad2b6272e8090ebc9c7cf36a490975d

Observation 18c493fd-f9b1-4e8b-9689-78de7603398e · outbound

This paper cites an unresolved cited work.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Unresolved cited work

Reference 57

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T11:56:34.382001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.324022Z digest=sha256:3f3e1ffbbd63bfa10d363dbd70066f3773c8bb015dcc07b44d69f9e7017ed652

Observation 64e7e780-3939-43a2-953a-1324e6ab1c40 · outbound

This paper cites Language Models are Open Knowledge Graphs.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Language Models are Open Knowledge Graphs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.327247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.327247Z digest=sha256:6ec1714bc2e528e0907b3a2abfc3583763abaf5d4e4a03ec5598d722a04d876e

Observation 60d4733b-45b1-4bef-880f-000b5a15117c · outbound

This paper cites Can AI Be as Creative as Humans?.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Can AI Be as Creative as Humans?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.330717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.330717Z digest=sha256:9090272d069af8cc33e63b34f384356666d6752aef8aa1edb8abfea3a2fd8225

Observation 5d323f19-07f3-41a9-9f95-676aaa7aae99 · outbound

This paper cites Chi, Quoc V.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Chi, Quoc V

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.830980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.334226Z digest=sha256:20e358cf4604c2657660b8cfd6418e97ca2b91108441f80f06a777011591763f

Observation 40608028-8344-4c9d-bbfe-42717f898ea3 · outbound

This paper cites Estimating the probabilities of rare outputs in language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the probabilities of rare outputs in language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.822545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.337670Z digest=sha256:1b9a4acd8cb19cf62745957206ad787e488dbc61537b73e7dd1a980a07ac4177

Observation 8f4f37ea-0c4b-4110-a534-f91081bc8a5e · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Qwen2.5 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.340659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.340659Z digest=sha256:48b749ba9e039820bf326cfdd536bc80a6d4e9cbc9bed582b1be56bea6b59195

Observation 893d72c8-1572-424b-9370-94f79a7c44b2 · outbound

This paper cites A careful examination of large language model performance on grade school arithmetic.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A careful examination of large language model performance on grade school arithmetic

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.812963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.343378Z digest=sha256:87c931f455e6a41fdd16b3ab58f17d63bf21d491c7acc2d8f9d80ddd12f4e979

Observation 88fcaddc-d668-4a60-87fd-a06a821f89a5 · outbound

This paper cites Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:56:34.413661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.346706Z digest=sha256:df4ac4631d1e49baadf0cf601d28a8aa44f12165b973549cbbfd234313ae9abd

Observation 29d200ca-a50b-460f-8047-70ddb01d209f · outbound

This paper cites Test your knowledge of mathematical theorems by listing 20 theorem names, separated by commas.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Test your knowledge of mathematical theorems by listing 20 theorem names, separated by commas

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:56:34.803679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:56:34.350227Z digest=sha256:118e00440b2dd1eb874a7b0080db5cd882f196e27a8669d09f58c14333c75e80

Pith citing papers

Observation 4f39711d-05df-4367-bdf0-2987ca8ae6db · inbound

UCS: Estimating Unseen Coverage for Improved In-Context Learning cites this paper.

UCS: Estimating Unseen Coverage for Improved In-Context Learning Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:01.283344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:40:18.133637Z digest=sha256:0a65ce01d16a0166120965bb149b9fce89fbcf877272ba417485ac7a364667b6