Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:38:22.316432Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2504.15208.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:38:22.316432Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:39.503440Z
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7b3cf814-384d-4a92-814d-15bb958031fd · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96795756-2584-4300-b416-79dfe1933a1a · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Understanding prompt engineering may not require rethinking generalization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ab4c9f-cb2b-47f6-a73a-75b6da82d5ab · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Stronger generalization bounds for deep nets via a compression approach
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d84c9bb0-e85f-43ad-9058-0150d8bb08f8 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Explaining neural scaling laws
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 44c11671-9354-4433-919c-7ef63e26aaf6 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Chinchilla Scaling: A replication attempt
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a2d3eb-d209-47d2-86d5-37416d71743a · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Pythia: A suite for analyzing large language models across training and scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db98836-ce42-4401-b627-50ca42b6e204 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The description length of deep learning models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc3d543-7ad7-4e03-b448-69a26a1e5dda · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The tradeoffs of large scale learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f885886-6807-417f-8db1-acfa8a94417e · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Bias/variance is not the same as approximation/estimation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 363fe238-5f9d-4b71-bbdb-85232bf6262d · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Language Models are Few-Shot Learners
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5bf52f-1a7e-4347-9924-b4fc0865e578 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Cleaning large correlation matrices: tools from random matrix theory
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e75924da-6c54-4e14-a1a7-e5a97464db7e · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Pac-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b5f6c3b-3702-4176-a2a1-fd0871ab1ea8 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Quip: 2-bit quantization of large language models with guarantees
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67746957-f5e9-49d7-aa8f-18f31f8fdedb · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale A unified recipe for deriving (time-uniform) pac-bayes bounds
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2643b37-8139-4f60-a096-c99a2ce3f8c8 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd27181-e3d3-4054-89ad-c06ff780b013 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Present position and potential developments: Some personal views statistical theory the prequential approach
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2df1ef4e-afb7-485e-adcc-315b62827b67 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4744cd00-f691-4daa-8229-1e453028253d · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Scalable log determinants for gaussian process kernel learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3bcf3f8-314a-49d7-a0c8-52de742ad0f5 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aceb9ed-75e5-4346-9c47-4b8479e12402 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b464350-61d8-4ed9-81a7-9de9a51e670f · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Entropic trace estimates for log determinants
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3cf335c6-2e66-42d3-af3a-29241c302919 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Gptq: Accurate post-training quantization for generative pre-trained transformers, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03c8329-df62-436e-bfa9-4c27f9f7def1 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Freedman
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 83ef1ba5-b86f-45d7-ab06-dbba6a2fd77f · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 415087f9-5cab-4215-b5e3-f4f8967eaaab · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale An investigation into neural net optimization via hessian eigenvalue density
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849ec5a8-8b01-49fa-b315-017752f2caae · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Mlrg deep curvature: An open-source package to analyse and visualise neural network curvature and loss surface
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fa6e4084-21b5-4636-9fad-777f63fce707 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The deep learning limit: are negative neural network eigenvalues just noise? In ICML 2019 workshop on theoretical physics for deep learning, 2019
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f520a299-d2ca-42ef-8d59-032607d860fc · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Learning rates as a function of batch size: A random matrix theory approach to neural network training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39d7c201-f8e2-4c71-ad76-0e34123ee9fb · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Large language models are zero-shot time series forecasters
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b61e3b02-788b-44d7-a5e3-f0f1afe91f81 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Stork, and Gregory J
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b86b374e-5705-4024-88c0-4fe718beb772 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Scaling Laws for Autoregressive Generative Modeling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2072898-fc0c-4244-8adb-8c9d99a1cc5b · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Training Compute-Optimal Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a493637-6881-4bd5-9fd9-d11bc7b95fff · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale LoRA: Low-Rank Adaptation of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1211ec-96f3-4eca-9d80-0f72d1221657 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Accurate post training quantization with small calibration sets
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb16290f-2d34-40c6-9a84-88fcc7cf36f8 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Scaling Laws for Neural Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e07eec-742c-4668-95ab-aacd7d5c7248 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale A device for quantizing, grouping, and coding amplitude-modulated pulses
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 354aaba0-32d5-431d-82ff-90a55961e9b4 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Transformers as algorithms: Generalization and stability in in-context learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1ea776a1-dd85-46c6-adef-1c2e2b61731d · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Lee, Song Han, Tri Dao, and Tianle Cai
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f484ace4-4f99-470e-a4a4-5e549bebc658 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Pac-bayes compression bounds so tight that they can explain generalization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a50627d6-909e-4696-9d89-7346b2e06f0a · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1c01941a-7b69-47d8-96ff-13dc202e805d · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6934890-b0ec-4c67-b391-fd75ad7e994a · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The era of 1-bit llms: All large language models are in 1.58 bits, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36286e9a-b7c7-4cf1-88dc-0de0eb1cec8e · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Empirical Bernstein Bounds and Sample Variance Penalization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1eeec9-0a2e-4057-aa1c-8fd4d1e1bb56 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Two inequalities implied by unique decipherability
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1b2f8dfc-9c07-44a1-b66e-f7a6ac24da85 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale The lanczos and conjugate gradient algorithms in finite precision arithmetic
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c020fb14-7aec-4912-86c2-f73e00e9fea3 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Scaling data-constrained language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d810169b-38a5-4842-bb70-6a597dca3954 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Up or down? adaptive rounding for post-training quantization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e00a7dd6-0e83-4b18-9781-a981cefa2c4a · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Gpt-4 technical report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf771928-32e8-4534-8178-3f79fa0ff8aa · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a77ed7-a18e-4dbc-af46-f3d4fb3a4dbc · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Mapping language models to grounded conceptual spaces
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e7743e77-bd73-464e-99cf-a243f138b066 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Fast exact multiplication by the hessian
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7945c6-c682-4553-baa5-aaf78f95b3eb · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale CoLA: Exploiting Compositional Structure for Automatic and Efficient Numerical Linear Algebra
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e7de62fb-6cc7-4a7e-a39f-d1ef9c980ce5 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Universal coding, information, prediction, and estimation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53d00811-e41c-48e7-8d16-227f9b68dfbd · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Improved bounds on sample size for implicit matrix trace estimators
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 46d5a63c-6a49-4806-b51d-e08858ea860b · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b090e3-87f3-4ffb-aaf3-3f6b63d8aec0 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Understanding machine learning: From theory to algorithms
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ea09fc-0164-4d79-a124-ca12a33bf2fd · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale A formal theory of inductive inference
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 855d97b9-8de3-40e6-a490-826520283181 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Trinh, Yuhuai Wu, Quoc V
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 12c15269-5b54-4fdb-9dbd-0cbb7e78a484 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Quip: Even better llm quantization with hadamard incoherence and lattice codebooks, 2024
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4385a41d-0cc7-471a-94b5-f01162beb711 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Fast estimation of tr(f(a)) via stochastic lanczos quadrature
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bc52a9-b5c5-4103-a42c-762db9e76861 · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Étude critique de la notion de collectif
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b4d2b1-a120-4fca-a0f1-c81580b8f8cc · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Information-Theoretic Probing with Minimum Description Length
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a49b61a-c3c2-46dd-a9a7-c8c2c44b479b · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa136ff-7aef-4a09-a51d-bd80912c5e2e · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Measuring Information Transfer in Neural Networks
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b67858b-2ad6-45d4-b272-117ad8f13fff · outbound
Compute-Optimal LLMs Provably Generalize Better With Scale Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0430e10-3c59-4534-997b-f8613924f1d5 · inbound
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search Compute-Optimal LLMs Provably Generalize Better With Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.