Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:16:10.422836Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 3 inbound Pith citation observations for arXiv:2505.10559.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:16:10.422836Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:05.518604Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T07:00:43.273983Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d212e579-4736-4103-b4d3-8542c19ba52b · outbound
Neural Thermodynamic Laws for Large Language Model Training Bayesian learning via stochastic gradient langevin dynamics
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a950008-b592-48e9-b839-c497e6d402ee · outbound
Neural Thermodynamic Laws for Large Language Model Training Thermodynamics-inspired explanations of artificial intelli- gence
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eb823f02-bcc2-435c-b0db-c5cca54d8faa · outbound
Neural Thermodynamic Laws for Large Language Model Training Statistical mechanics of learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4b77da36-8744-449b-9ffa-0d35359d4f31 · outbound
Neural Thermodynamic Laws for Large Language Model Training Statistical mechanics of deep learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c0f021a9-d8f0-4eab-8225-95664e1e90bc · outbound
Neural Thermodynamic Laws for Large Language Model Training Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e2a7b6-6bce-48d9-8fae-5acc1ef8208e · outbound
Neural Thermodynamic Laws for Large Language Model Training How noise affects the Hessian spectrum in overparameterized neural networks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5705197c-17c9-408b-b4a5-fb1d4f50b76b · outbound
Neural Thermodynamic Laws for Large Language Model Training FOCUS: First Order Concentrated Updating Scheme
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1eea92-38df-4956-89e6-2fc88a336a71 · outbound
Neural Thermodynamic Laws for Large Language Model Training MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be573704-5082-4bfe-8460-94297b5b5030 · outbound
Neural Thermodynamic Laws for Large Language Model Training Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58f5b92-0d8c-46dd-a841-f45d69571120 · outbound
Neural Thermodynamic Laws for Large Language Model Training Scaling laws and compute-optimal training beyond fixed training durations
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 54d121ff-1a49-4fc7-bda8-ba7003fa8064 · outbound
Neural Thermodynamic Laws for Large Language Model Training Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9dbac8-217c-4ea7-8ed3-ea12abddc571 · outbound
Neural Thermodynamic Laws for Large Language Model Training A multi-power law for loss curve prediction across learning rate schedules
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation adb2480a-b973-45ce-8c44-4d3486f0c5e2 · outbound
Neural Thermodynamic Laws for Large Language Model Training modded-nanogpt
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e10bec04-7f0f-4a15-9eab-f4feff0a122d · outbound
Neural Thermodynamic Laws for Large Language Model Training Implicit Gradient Regularization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 975e5bbb-1674-458d-926f-1134ba17a91b · outbound
Neural Thermodynamic Laws for Large Language Model Training The limiting dynamics of sgd: Modified loss, phase-space oscillations, and anomalous diffusion
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3fa17f3c-0219-4868-9251-f6ded64f650d · outbound
Neural Thermodynamic Laws for Large Language Model Training Stochastic collapse: How gra- dient noise attracts sgd dynamics towards simpler subnetworks.Advances in Neural Information Processing Systems, 36:35027–35063, 2023
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 710f4bf1-39bc-443b-93ce-34b437911117 · outbound
Neural Thermodynamic Laws for Large Language Model Training Stochastic gradient descent as approximate bayesian inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e756cd0-1ae7-49aa-b88d-f3b250cec91c · outbound
Neural Thermodynamic Laws for Large Language Model Training A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89849d2-59be-42e7-80ce-2cfc42587224 · outbound
Neural Thermodynamic Laws for Large Language Model Training Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dbec16f3-8046-48dd-9e77-15f4b5af7ac0 · outbound
Neural Thermodynamic Laws for Large Language Model Training Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d8dc6a4-c009-475a-a7e6-526b7fd195d4 · outbound
Neural Thermodynamic Laws for Large Language Model Training The implicit bias for adaptive optimization algorithms on homogeneous neural networks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5bf4fd29-7f62-4325-bc6a-6971e63bb0e1 · outbound
Neural Thermodynamic Laws for Large Language Model Training An overview of condensation phenomenon in deep learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30cc1f6-f2d3-4bec-8fef-f587ab01b983 · outbound
Neural Thermodynamic Laws for Large Language Model Training Loss surfaces, mode connectivity, and fast ensembling of dnns
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6fee98e-5454-4dfd-98ee-63807f5f9ac9 · outbound
Neural Thermodynamic Laws for Large Language Model Training Linear mode connectivity and the lottery ticket hypothesis
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14003317-7996-4b1c-aa40-4bcd29f84b95 · outbound
Neural Thermodynamic Laws for Large Language Model Training SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c19445-75c5-4c60-b43b-ed7169e8d14d · outbound
Neural Thermodynamic Laws for Large Language Model Training Cyclical learning rates for training neural networks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cae9e569-d9fb-44df-8f14-d605328e46c4 · outbound
Neural Thermodynamic Laws for Large Language Model Training Attention is all you need
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389cbea2-8c0b-46ca-a545-57a66aebce39 · outbound
Neural Thermodynamic Laws for Large Language Model Training The information bottleneck method
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf45d999-1565-418a-abed-41956b681423 · outbound
Neural Thermodynamic Laws for Large Language Model Training Entropy-sgd: Biasing gradient descent into wide valleys
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0221e432-621f-4859-9a46-591124545197 · outbound
Neural Thermodynamic Laws for Large Language Model Training A learning algorithm for boltzmann machines
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98bdeea6-d6e2-4d6d-96de-4eb885ceffae · outbound
Neural Thermodynamic Laws for Large Language Model Training Hopfield Networks is All You Need
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c2a8453-6e38-4eb4-804f-834e09498e3a · outbound
Neural Thermodynamic Laws for Large Language Model Training Deep unsuper- vised learning using nonequilibrium thermodynamics
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cca370-659f-4db5-9d77-e0379a008b63 · outbound
Neural Thermodynamic Laws for Large Language Model Training Score-Based Generative Modeling through Stochastic Differential Equations
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8430a16a-5d30-48c7-a45d-43a0ab8ab73c · outbound
Neural Thermodynamic Laws for Large Language Model Training edge of stability
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 52d8cbf4-a455-4428-b670-c19fcb8309d5 · inbound
Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model Neural Thermodynamic Laws for Large Language Model Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88f9cc7-7416-4d8d-b8e8-91e37a2868d3 · inbound
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Neural Thermodynamic Laws for Large Language Model Training
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4617356d-24ea-480a-be5f-e57e497bff5e · inbound
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Neural Thermodynamic Laws for Large Language Model Training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.