Pith. sign in

Paper Citation Record · LEDGER

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2506.09099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09099 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:55.962152Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:02:06.788056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:46:25.975143Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8161e917-cecd-4e3a-a2a4-1c889c3a655e · outbound

This paper cites GPT-4 Technical Report.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.911995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.911995Z digest=sha256:39bd9991eeae9e3ab04e7ad24f965ea3b4c317cdff2a4b101d42dafe347d0fdd

Observation 57c1af3c-30f3-4ba1-8fa7-3fe26dac7beb · outbound

This paper cites 9, 2024); accessed May 19,.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers 9, 2024); accessed May 19,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:56.167851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:04:55.925247Z digest=sha256:bfba6f152a86f795743e4396b0c7d49182de5a7e35dc590084feaa95d4cb5fa1

Observation 04fcb939-cc24-4f4e-89c4-98b4e2b86369 · outbound

This paper cites doi: 10.1016/j.neunet.2024.106550.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers doi: 10.1016/j.neunet.2024.106550

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.932743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.932743Z digest=sha256:175635285b8810958dbb12fc63e8231ece37f28c397199b49e7de63aceee5086

Observation d55ccfad-9b27-495b-b0a9-0810362e5e9e · outbound

This paper cites Training language models to follow instructions with human feedback.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.936614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.936614Z digest=sha256:3fd8ca79300bb4b78a6d100f5e66db8a4081048a3a96f58babcae007695191bb

Observation e040c504-c48e-47f9-8502-6c11da144c51 · outbound

This paper cites Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.940338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.940338Z digest=sha256:1b44aa6a1454668ace01f4bfdcf0685d2a68a19c4cb6c2f75da7bda2f7adda87

Observation 170e1a5b-513d-4023-b290-12eee9943dc2 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.944197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.944197Z digest=sha256:873441e78b724825c56812a051093a6b500d1312520f8a53ec3392ef30bac1e6

Observation d414a4b0-7877-4fb0-bf7d-2931d92c6792 · outbound

This paper cites Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.951254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.951254Z digest=sha256:4655ff84e2d456bd885bb507056d7d4bd9ba0bbe4c370f4bc16d0fcd9eb7cd7d

Observation 43e43de8-b55d-4d46-8ce2-d425768a8620 · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.954824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.954824Z digest=sha256:818db7c07d181145a831aaadae2bbe867b4d89aac9fe5d5798587fa44ff27e78

Observation 456e782f-fb2c-42de-9a69-8828c499f5c4 · outbound

This paper cites Exploring Memorization in Fine-tuned Language Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Exploring Memorization in Fine-tuned Language Models

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:04:55.958433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.958433Z digest=sha256:db1addaf86e903b542ae670e638b62674b34096511aeb6bcad4c3de94ba6a67d

Observation 48de51b3-0a53-43ee-91df-0e18c07519cb · outbound

This paper cites Goldilocks zone.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Goldilocks zone

Reference 14

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:04:56.157226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:04:55.962152Z digest=sha256:9c3a88b9274fb8ba6a7757bac39993381c0770108bcd0cad8b8bb81932607232

Observation 6083ae0c-f92c-4cd0-93e5-190a11b12064 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Gemini: A Family of Highly Capable Multimodal Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.948019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.948019Z digest=sha256:7767fe46500b8c2c132822a2ce85e9b92ccf8b53261a269c235d2a693ef3c170

Observation f86a55e6-0981-4687-a4ba-da29f92be046 · outbound

This paper cites Towards Understanding Grokking: An Effective Theory of Representation Learning.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers Towards Understanding Grokking: An Effective Theory of Representation Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.929001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.929001Z digest=sha256:481e382819f338d1540ccbb546b2ad8a9983f9ad9beb712d1d4819f1c951b08c

Observation 7d4bf7da-300f-490b-b06e-29a15f30c65a · outbound

This paper cites GPT-4o System Card.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.921232Z digest=sha256:0c1b9d2535117920fcfdec23ece1585965feb97d9486ddec5a65928f5936a308

Observation a0b355b9-aa0d-4b4d-9c9b-0e98309b49f6 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:55.916932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:55.916932Z digest=sha256:d462c0eb14d5116e02dba90c46f7660d0552c8564fc8f6342c88a30b2882e613

Pith citing papers

Observation 6a6da3b7-5808-4654-9ae7-d0e6edac71f7 · inbound

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities cites this paper.

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:25.982733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:02:06.788056Z digest=sha256:988b6b87a8135e1564854e59876c8ced8feb9a0ba96a027b643c3bd4f3ce6b93