Pith. sign in

Paper Citation Record · LEDGER

Compute-Optimal LLMs Provably Generalize Better With Scale

As of 19 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2504.15208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15208 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:38:22.316432Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:39.503440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b3cf814-384d-4a92-814d-15bb958031fd · outbound

This paper cites write newline.

Compute-Optimal LLMs Provably Generalize Better With Scale write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.075061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.075061Z digest=sha256:b6edc02208639eb7a3cca8cf285ecf13812c7f1197500316db89860663f7c8b6

Observation 96795756-2584-4300-b416-79dfe1933a1a · outbound

This paper cites Understanding prompt engineering may not require rethinking generalization.

Compute-Optimal LLMs Provably Generalize Better With Scale Understanding prompt engineering may not require rethinking generalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.079896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.079896Z digest=sha256:f4044ecdb886b2192fbfe159075b26d1fcaaa309c1ec1b4ec84f9ae02734fbb4

Observation 37ab4c9f-cb2b-47f6-a73a-75b6da82d5ab · outbound

This paper cites Stronger generalization bounds for deep nets via a compression approach.

Compute-Optimal LLMs Provably Generalize Better With Scale Stronger generalization bounds for deep nets via a compression approach

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.977158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.085392Z digest=sha256:3a8b376a72bf12ae8077c2e8b8763e11f5f4ee1c8d1e4199a153d088393bc16f

Observation d84c9bb0-e85f-43ad-9058-0150d8bb08f8 · outbound

This paper cites Explaining neural scaling laws.

Compute-Optimal LLMs Provably Generalize Better With Scale Explaining neural scaling laws

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.965870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.088934Z digest=sha256:2b8ade3c714ee2f0edd8fcf90d86d74e4c5fbf70cbcf9ad626088f0452f7e6d1

Observation 44c11671-9354-4433-919c-7ef63e26aaf6 · outbound

This paper cites Chinchilla Scaling: A replication attempt.

Compute-Optimal LLMs Provably Generalize Better With Scale Chinchilla Scaling: A replication attempt

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.092818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.092818Z digest=sha256:e0421123a288360b08bde47312eb426356ad06d51471b585747f3af77be7f766

Observation 94a2d3eb-d209-47d2-86d5-37416d71743a · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Compute-Optimal LLMs Provably Generalize Better With Scale Pythia: A suite for analyzing large language models across training and scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.096779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.096779Z digest=sha256:c77517ef66b06a9aceb3e12b1941dd9987de98915f39f1e613bc3b19e6da8973

Observation 0db98836-ce42-4401-b627-50ca42b6e204 · outbound

This paper cites The description length of deep learning models.

Compute-Optimal LLMs Provably Generalize Better With Scale The description length of deep learning models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.100555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.100555Z digest=sha256:2834e547ae94068c313e0ce4eb0f7a6c2df183a10f9cd82d39f9045064bc69e9

Observation 9bc3d543-7ad7-4e03-b448-69a26a1e5dda · outbound

This paper cites The tradeoffs of large scale learning.

Compute-Optimal LLMs Provably Generalize Better With Scale The tradeoffs of large scale learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.104403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.104403Z digest=sha256:a7ba927f2308025efc8ca8fff553b3052a3482eef84383fb53d30560b162f3e5

Observation 6f885886-6807-417f-8db1-acfa8a94417e · outbound

This paper cites Bias/variance is not the same as approximation/estimation.

Compute-Optimal LLMs Provably Generalize Better With Scale Bias/variance is not the same as approximation/estimation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.933298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.107937Z digest=sha256:93cd6f5ec752ff676944e4748e2a6bf8daf2b8ac313ca5e396fac081fa78f27b

Observation 363fe238-5f9d-4b71-bbdb-85232bf6262d · outbound

This paper cites Language Models are Few-Shot Learners.

Compute-Optimal LLMs Provably Generalize Better With Scale Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.111193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.111193Z digest=sha256:17a4688651bd02d753bf2ace2230ed7750daddafd4a6e131dfbd7c3f6e6c9540

Observation eb5bf52f-1a7e-4347-9924-b4fc0865e578 · outbound

This paper cites Cleaning large correlation matrices: tools from random matrix theory.

Compute-Optimal LLMs Provably Generalize Better With Scale Cleaning large correlation matrices: tools from random matrix theory

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.922711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.115311Z digest=sha256:ba6b003a3fa73831978831e19d3dfcf2cd1bce34b03efaf8cb5158f2ba6e96a0

Observation e75924da-6c54-4e14-a1a7-e5a97464db7e · outbound

This paper cites Pac-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning.

Compute-Optimal LLMs Provably Generalize Better With Scale Pac-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.118996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.118996Z digest=sha256:b5f62a974d49064ba00ec004ca47fc681a0e8ef0a043edbba60faed92e52a800

Observation 6b5f6c3b-3702-4176-a2a1-fd0871ab1ea8 · outbound

This paper cites Quip: 2-bit quantization of large language models with guarantees.

Compute-Optimal LLMs Provably Generalize Better With Scale Quip: 2-bit quantization of large language models with guarantees

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.122850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.122850Z digest=sha256:a3e489b0047f1d1cf73f868fd11dab7c657b4734a06ed437783aa1066b8e737e

Observation 67746957-f5e9-49d7-aa8f-18f31f8fdedb · outbound

This paper cites A unified recipe for deriving (time-uniform) pac-bayes bounds.

Compute-Optimal LLMs Provably Generalize Better With Scale A unified recipe for deriving (time-uniform) pac-bayes bounds

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.905886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.126372Z digest=sha256:226e7495f3a22289951884860a7302713dcb384a8b18626f1edfeb9641c582b6

Observation d2643b37-8139-4f60-a096-c99a2ce3f8c8 · outbound

This paper cites an unresolved cited work.

Compute-Optimal LLMs Provably Generalize Better With Scale Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.129893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.129893Z digest=sha256:fce6082d5254fa3f81edbf643a3251a84af9fa88e2a948e15ea241889fe82d98

Observation 0fd27181-e3d3-4054-89ad-c06ff780b013 · outbound

This paper cites Present position and potential developments: Some personal views statistical theory the prequential approach.

Compute-Optimal LLMs Provably Generalize Better With Scale Present position and potential developments: Some personal views statistical theory the prequential approach

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.895212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.133517Z digest=sha256:66a38e65a6998824b1adcc36d4872d4ee26bd264a2a0e9fc4290c10774e33231

Observation 2df1ef4e-afb7-485e-adcc-315b62827b67 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

Compute-Optimal LLMs Provably Generalize Better With Scale LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.137161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.137161Z digest=sha256:9ca034d126b6b77fb41226681409c351039f5fb1f2c84a1e410fcb5c2e3ea663

Observation 4744cd00-f691-4daa-8229-1e453028253d · outbound

This paper cites Scalable log determinants for gaussian process kernel learning.

Compute-Optimal LLMs Provably Generalize Better With Scale Scalable log determinants for gaussian process kernel learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.883274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.141052Z digest=sha256:ee877cddc79cfa0cbea8c6e4594a2dc825c07e760a2bb9517972bbd11af1ae04

Observation f3bcf3f8-314a-49d7-a0c8-52de742ad0f5 · outbound

This paper cites The Llama 3 Herd of Models.

Compute-Optimal LLMs Provably Generalize Better With Scale The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.144884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.144884Z digest=sha256:83401c5f59f33f08c3fdc709b13200e6069b05b81285361b2113014a131d8df6

Observation 9aceb9ed-75e5-4346-9c47-4b8479e12402 · outbound

This paper cites Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data.

Compute-Optimal LLMs Provably Generalize Better With Scale Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.148737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.148737Z digest=sha256:487cd1997741fac6a5ceda7c14b8181194ade7c6da9709dc9b77cc655037293c

Observation 9b464350-61d8-4ed9-81a7-9de9a51e670f · outbound

This paper cites Entropic trace estimates for log determinants.

Compute-Optimal LLMs Provably Generalize Better With Scale Entropic trace estimates for log determinants

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.870952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.152776Z digest=sha256:aa6f642ce578989d852a4ff3e8602ae3e7417f5e2ea6154a0e4ed3fd3e4db423

Observation 3cf335c6-2e66-42d3-af3a-29241c302919 · outbound

This paper cites Gptq: Accurate post-training quantization for generative pre-trained transformers, 2023.

Compute-Optimal LLMs Provably Generalize Better With Scale Gptq: Accurate post-training quantization for generative pre-trained transformers, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.156151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.156151Z digest=sha256:26e61deaec6b12584cd5b72d17224a0519ffb5839cd873725dd302ee52a873eb

Observation c03c8329-df62-436e-bfa9-4c27f9f7def1 · outbound

This paper cites Freedman.

Compute-Optimal LLMs Provably Generalize Better With Scale Freedman

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.852798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.159875Z digest=sha256:f9d698a9f3b0dffe874d5be02b62790587a7ca17d5a7ae5ba23cae3fb11769dc

Observation 83ef1ba5-b86f-45d7-ab06-dbba6a2fd77f · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Compute-Optimal LLMs Provably Generalize Better With Scale The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.163291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.163291Z digest=sha256:89010ef88db3bb1e3c758a3391686ed4655dd61d111e0a99aa7e67d6259a368f

Observation 415087f9-5cab-4215-b5e3-f4f8967eaaab · outbound

This paper cites An investigation into neural net optimization via hessian eigenvalue density.

Compute-Optimal LLMs Provably Generalize Better With Scale An investigation into neural net optimization via hessian eigenvalue density

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.167147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.167147Z digest=sha256:512c2c1f4689071c8e0f5d7958a9e2066c63920e718afca83deaf184a3170208

Observation 849ec5a8-8b01-49fa-b315-017752f2caae · outbound

This paper cites Mlrg deep curvature: An open-source package to analyse and visualise neural network curvature and loss surface.

Compute-Optimal LLMs Provably Generalize Better With Scale Mlrg deep curvature: An open-source package to analyse and visualise neural network curvature and loss surface

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.835426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.170829Z digest=sha256:4ebbd8728c0cb3413a0c0651b9e65fd7c5c6ec5cc6d523d77963b7da5b7eff71

Observation fa6e4084-21b5-4636-9fad-777f63fce707 · outbound

This paper cites The deep learning limit: are negative neural network eigenvalues just noise? In ICML 2019 workshop on theoretical physics for deep learning, 2019.

Compute-Optimal LLMs Provably Generalize Better With Scale The deep learning limit: are negative neural network eigenvalues just noise? In ICML 2019 workshop on theoretical physics for deep learning, 2019

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.824592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.174611Z digest=sha256:3b9e5ee22d1bdaadd6b0e9863073622d4bb5291103698d57f62ae544db3dac01

Observation f520a299-d2ca-42ef-8d59-032607d860fc · outbound

This paper cites Learning rates as a function of batch size: A random matrix theory approach to neural network training.

Compute-Optimal LLMs Provably Generalize Better With Scale Learning rates as a function of batch size: A random matrix theory approach to neural network training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.814018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.177907Z digest=sha256:3ae5427627b3db3a9970d54c467c53e7bcba97ec9c6a37977de70d06a36c06cf

Observation 39d7c201-f8e2-4c71-ad76-0e34123ee9fb · outbound

This paper cites Large language models are zero-shot time series forecasters.

Compute-Optimal LLMs Provably Generalize Better With Scale Large language models are zero-shot time series forecasters

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.803294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.181619Z digest=sha256:f80408838f8855f1c5a746e084cdf3c74566442986320f4ccff74fe948956659

Observation b61e3b02-788b-44d7-a5e3-f0f1afe91f81 · outbound

This paper cites Stork, and Gregory J.

Compute-Optimal LLMs Provably Generalize Better With Scale Stork, and Gregory J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.793136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.185340Z digest=sha256:0b867933c213366c55b2c3030a74e279914e7e80fc534ed8d354d29b1f83e029

Observation b86b374e-5705-4024-88c0-4fe718beb772 · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Compute-Optimal LLMs Provably Generalize Better With Scale Scaling Laws for Autoregressive Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.188868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.188868Z digest=sha256:81fb568a1ad517a5cfc10aa8a22fe4d1219998e23022ae992945bc0c9ec4db6f

Observation b2072898-fc0c-4244-8adb-8c9d99a1cc5b · outbound

This paper cites Training Compute-Optimal Large Language Models.

Compute-Optimal LLMs Provably Generalize Better With Scale Training Compute-Optimal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.192753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.192753Z digest=sha256:ec9dd649a65fa44a09d3dcd9dd90f4c9e4a0ff8e02d6c816ec26ddc243a2241a

Observation 4a493637-6881-4bd5-9fd9-d11bc7b95fff · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Compute-Optimal LLMs Provably Generalize Better With Scale LoRA: Low-Rank Adaptation of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.196353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.196353Z digest=sha256:61a1397893aed705fbe51f7b5851390448e0b836ba183765189d907e3974d2a0

Observation 0f1211ec-96f3-4eca-9d80-0f72d1221657 · outbound

This paper cites Accurate post training quantization with small calibration sets.

Compute-Optimal LLMs Provably Generalize Better With Scale Accurate post training quantization with small calibration sets

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.782355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.199958Z digest=sha256:71aa92016c6bc9d31f951d164b77219509a5141eb8dd1e11509cccdf6bb87e76

Observation bb16290f-2d34-40c6-9a84-88fcc7cf36f8 · outbound

This paper cites Scaling Laws for Neural Language Models.

Compute-Optimal LLMs Provably Generalize Better With Scale Scaling Laws for Neural Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.203753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.203753Z digest=sha256:f3a3e77e3274ec88d6558d969005ad60a352899a89e918d9b7c6d2fc26cfd168

Observation 21e07eec-742c-4668-95ab-aacd7d5c7248 · outbound

This paper cites A device for quantizing, grouping, and coding amplitude-modulated pulses.

Compute-Optimal LLMs Provably Generalize Better With Scale A device for quantizing, grouping, and coding amplitude-modulated pulses

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.770579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.207380Z digest=sha256:f8dbe8457a23fd5618617d09fde7275abd80c312ac449a2b7f9302b43c4a76e8

Observation 354aaba0-32d5-431d-82ff-90a55961e9b4 · outbound

This paper cites Transformers as algorithms: Generalization and stability in in-context learning.

Compute-Optimal LLMs Provably Generalize Better With Scale Transformers as algorithms: Generalization and stability in in-context learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.759739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.211093Z digest=sha256:d5baf186495e434ff2c6aae2f2b9f3fc03ef751d2dd97db6dda39f77ae13032b

Observation 1ea776a1-dd85-46c6-adef-1c2e2b61731d · outbound

This paper cites Lee, Song Han, Tri Dao, and Tianle Cai.

Compute-Optimal LLMs Provably Generalize Better With Scale Lee, Song Han, Tri Dao, and Tianle Cai

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.748533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.214717Z digest=sha256:4ef3f8eb898c931ac2f76893b76aa828c96245ec8b56eda93edb262aafe57346

Observation f484ace4-4f99-470e-a4a4-5e549bebc658 · outbound

This paper cites Pac-bayes compression bounds so tight that they can explain generalization.

Compute-Optimal LLMs Provably Generalize Better With Scale Pac-bayes compression bounds so tight that they can explain generalization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.738268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.218528Z digest=sha256:0cc1b89354ca2496a96e6cdcb10a7fd68ff5899f3a721607224d58c56bb1b3be

Observation a50627d6-909e-4696-9d89-7346b2e06f0a · outbound

This paper cites an unresolved cited work.

Compute-Optimal LLMs Provably Generalize Better With Scale Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:38:22.727805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.221895Z digest=sha256:1f601d5f33d810394bdde67506e7bc736ed44028d14fe809dc86a0791abd8bbd

Observation 1c01941a-7b69-47d8-96ff-13dc202e805d · outbound

This paper cites Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models.

Compute-Optimal LLMs Provably Generalize Better With Scale Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.229148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.229148Z digest=sha256:025376ee5601f613d34f8cb3a980fe86e7b1c7a55d65464de62645bf578ee1d3

Observation a6934890-b0ec-4c67-b391-fd75ad7e994a · outbound

This paper cites The era of 1-bit llms: All large language models are in 1.58 bits, 2024.

Compute-Optimal LLMs Provably Generalize Better With Scale The era of 1-bit llms: All large language models are in 1.58 bits, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.232959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.232959Z digest=sha256:0b5fe950b8617589bf647ea4d64a0ce1e9f913dc97407a9fb1868544410c67a2

Observation 36286e9a-b7c7-4cf1-88dc-0de0eb1cec8e · outbound

This paper cites Empirical Bernstein Bounds and Sample Variance Penalization.

Compute-Optimal LLMs Provably Generalize Better With Scale Empirical Bernstein Bounds and Sample Variance Penalization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.236301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.236301Z digest=sha256:7948e594e884444e76dae3b929fcc00950476f1ace79d113906982a3b33d71ea

Observation 1c1eeec9-0a2e-4057-aa1c-8fd4d1e1bb56 · outbound

This paper cites Two inequalities implied by unique decipherability.

Compute-Optimal LLMs Provably Generalize Better With Scale Two inequalities implied by unique decipherability

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.709518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.240120Z digest=sha256:660d17b206b94130abe2f790e4efa347720ae5cff4a4b776757fd14491433eec

Observation 1b2f8dfc-9c07-44a1-b66e-f7a6ac24da85 · outbound

This paper cites The lanczos and conjugate gradient algorithms in finite precision arithmetic.

Compute-Optimal LLMs Provably Generalize Better With Scale The lanczos and conjugate gradient algorithms in finite precision arithmetic

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.698534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.243641Z digest=sha256:a0a0e3bcc1f4dded15f65c18793467a87cf1fba752285f44f71ccaa0cdff0898

Observation c020fb14-7aec-4912-86c2-f73e00e9fea3 · outbound

This paper cites Scaling data-constrained language models.

Compute-Optimal LLMs Provably Generalize Better With Scale Scaling data-constrained language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.247008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.247008Z digest=sha256:f5fe32d023829ae75fee2a037e296c02ba7ff4fe3ca6c4d37bca1762c46beee8

Observation d810169b-38a5-4842-bb70-6a597dca3954 · outbound

This paper cites Up or down? adaptive rounding for post-training quantization.

Compute-Optimal LLMs Provably Generalize Better With Scale Up or down? adaptive rounding for post-training quantization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.680484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.250360Z digest=sha256:e5aef78e8f31314cc277e133f352ee2f501e474a30c2b21821e0af54126cbc67

Observation e00a7dd6-0e83-4b18-9781-a981cefa2c4a · outbound

This paper cites Gpt-4 technical report.

Compute-Optimal LLMs Provably Generalize Better With Scale Gpt-4 technical report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.254259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.254259Z digest=sha256:6e1e8f26c92be0af8e72098ea1a5d97bf373d00ddeed7094f3a32c7e93bdaac4

Observation cf771928-32e8-4534-8178-3f79fa0ff8aa · outbound

This paper cites Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians.

Compute-Optimal LLMs Provably Generalize Better With Scale Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.257738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.257738Z digest=sha256:3e382b8ec8c54f2015ea8d5ea8f62c0713ca852943127d2596650b5cee7a779c

Observation 95a77ed7-a18e-4dbc-af46-f3d4fb3a4dbc · outbound

This paper cites Mapping language models to grounded conceptual spaces.

Compute-Optimal LLMs Provably Generalize Better With Scale Mapping language models to grounded conceptual spaces

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.661537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.261846Z digest=sha256:5cb6e4b9cf492755858aee54077825b238d2425c036501268189e60cdf0418a3

Observation e7743e77-bd73-464e-99cf-a243f138b066 · outbound

This paper cites Fast exact multiplication by the hessian.

Compute-Optimal LLMs Provably Generalize Better With Scale Fast exact multiplication by the hessian

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.265462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.265462Z digest=sha256:3631fa0a4bc8dbcbccb1e3161d5d41f847656e95caecf3953a7a34d66d64bfd9

Observation af7945c6-c682-4553-baa5-aaf78f95b3eb · outbound

This paper cites CoLA: Exploiting Compositional Structure for Automatic and Efficient Numerical Linear Algebra.

Compute-Optimal LLMs Provably Generalize Better With Scale CoLA: Exploiting Compositional Structure for Automatic and Efficient Numerical Linear Algebra

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:38:22.410813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.269107Z digest=sha256:973740890751bac4f930ffc0f5de3a14ac33b88ad4e6079545a4bcc36739d535

Observation e7de62fb-6cc7-4a7e-a39f-d1ef9c980ce5 · outbound

This paper cites Universal coding, information, prediction, and estimation.

Compute-Optimal LLMs Provably Generalize Better With Scale Universal coding, information, prediction, and estimation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.644397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.272986Z digest=sha256:b077223689a2fb113a9e23725d08aff70e3a2fa85291f6f51b9e41dff6dc3e3c

Observation 53d00811-e41c-48e7-8d16-227f9b68dfbd · outbound

This paper cites Improved bounds on sample size for implicit matrix trace estimators.

Compute-Optimal LLMs Provably Generalize Better With Scale Improved bounds on sample size for implicit matrix trace estimators

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.633627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.276371Z digest=sha256:0ffe533f3213abbf938df466e722574d3476c72fb5bd532641b1073353e08d8d

Observation 46d5a63c-6a49-4806-b51d-e08858ea860b · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Compute-Optimal LLMs Provably Generalize Better With Scale Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.279976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.279976Z digest=sha256:f3268f916138d834a1823cb733ff19c93816654291f76c063f1303900a2a5c9a

Observation 59b090e3-87f3-4ffb-aaf3-3f6b63d8aec0 · outbound

This paper cites Understanding machine learning: From theory to algorithms.

Compute-Optimal LLMs Provably Generalize Better With Scale Understanding machine learning: From theory to algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.283762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.283762Z digest=sha256:24cef9b8e3e76f745434c119dc6483f244086b646b2b29364501713dafd1397f

Observation b0ea09fc-0164-4d79-a124-ca12a33bf2fd · outbound

This paper cites A formal theory of inductive inference.

Compute-Optimal LLMs Provably Generalize Better With Scale A formal theory of inductive inference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.615408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.287343Z digest=sha256:ef4016ae55713b84d4e0d45ae8b0144a79e8b134a9e9687162be1c454edf97be

Observation 855d97b9-8de3-40e6-a490-826520283181 · outbound

This paper cites Trinh, Yuhuai Wu, Quoc V.

Compute-Optimal LLMs Provably Generalize Better With Scale Trinh, Yuhuai Wu, Quoc V

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.604624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.291173Z digest=sha256:45a5919a24a959a3c83a0a0684a5d13aec93d7b5b5e4a21e81af4316648e4682

Observation 12c15269-5b54-4fdb-9dbd-0cbb7e78a484 · outbound

This paper cites Quip: Even better llm quantization with hadamard incoherence and lattice codebooks, 2024.

Compute-Optimal LLMs Provably Generalize Better With Scale Quip: Even better llm quantization with hadamard incoherence and lattice codebooks, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:38:22.593492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:38:22.294789Z digest=sha256:6354a4ea9e328698af4f81e01257ec421463f5be2d584813cd793e23f93a488e

Observation 4385a41d-0cc7-471a-94b5-f01162beb711 · outbound

This paper cites Fast estimation of tr(f(a)) via stochastic lanczos quadrature.

Compute-Optimal LLMs Provably Generalize Better With Scale Fast estimation of tr(f(a)) via stochastic lanczos quadrature

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.298302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.298302Z digest=sha256:f047a1dcd9b7ed48ba6dc9f144ca38ca060e611c9547045651b56dad1c7629f7

Observation 12bc52a9-b5c5-4103-a42c-762db9e76861 · outbound

This paper cites Étude critique de la notion de collectif.

Compute-Optimal LLMs Provably Generalize Better With Scale Étude critique de la notion de collectif

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.301721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.301721Z digest=sha256:378ee72ae2f061bd7f968dd9a934fce4bbe296ef9c8ce0c064c7e7eec5b621f8

Observation 05b4d2b1-a120-4fca-a0f1-c81580b8f8cc · outbound

This paper cites Information-Theoretic Probing with Minimum Description Length.

Compute-Optimal LLMs Provably Generalize Better With Scale Information-Theoretic Probing with Minimum Description Length

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.305478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.305478Z digest=sha256:576e5fb159a682ebc5da1652ce7a5d8825dd963882b81b565e861a24b7ff2cc0

Observation 6a49b61a-c3c2-46dd-a9a7-c8c2c44b479b · outbound

This paper cites ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers.

Compute-Optimal LLMs Provably Generalize Better With Scale ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.309181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.309181Z digest=sha256:699556c1620408a6d8c8d4b149adbd4875a1cdda4ec3aea0eea5ace6b20de50f

Observation 2fa136ff-7aef-4a09-a51d-bd80912c5e2e · outbound

This paper cites Measuring Information Transfer in Neural Networks.

Compute-Optimal LLMs Provably Generalize Better With Scale Measuring Information Transfer in Neural Networks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.312773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.312773Z digest=sha256:dad241d77528ee61c85e242ffdb3f7e0982734e8e2f0fc7dba84f4914c55c344

Observation 3b67858b-2ad6-45d4-b272-117ad8f13fff · outbound

This paper cites Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach.

Compute-Optimal LLMs Provably Generalize Better With Scale Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.316432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.316432Z digest=sha256:be713211593d69e61408dc8f61424652919df7a414360dc19af61a8bf109f997

Pith citing papers

Observation b0430e10-3c59-4534-997b-f8613924f1d5 · inbound

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search cites this paper.

A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search Compute-Optimal LLMs Provably Generalize Better With Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:39.503440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:39.503440Z digest=sha256:29eff627fe878fdcb204cb62ee3e3bc28fbf2eec4fb23f18d0831d9001406296