Pith. sign in

Paper Citation Record · LEDGER

Neural Thermodynamic Laws for Large Language Model Training

As of 18 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 3 inbound Pith citation observations for arXiv:2505.10559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10559 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:16:10.422836Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:05.518604Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T07:00:43.273983Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d212e579-4736-4103-b4d3-8542c19ba52b · outbound

This paper cites Bayesian learning via stochastic gradient langevin dynamics.

Neural Thermodynamic Laws for Large Language Model Training Bayesian learning via stochastic gradient langevin dynamics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.276576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.276576Z digest=sha256:3baa8686289958ebcf96e5749c2ce5e587060f3cb9e38da0660df7b1ede5d9be

Observation 6a950008-b592-48e9-b839-c497e6d402ee · outbound

This paper cites Thermodynamics-inspired explanations of artificial intelli- gence.

Neural Thermodynamic Laws for Large Language Model Training Thermodynamics-inspired explanations of artificial intelli- gence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.854755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.281206Z digest=sha256:1bcee09de21b07e9c4d64a6a6332238f0adb21914369782ef15f4d2513b2c316

Observation eb823f02-bcc2-435c-b0db-c5cca54d8faa · outbound

This paper cites Statistical mechanics of learning.

Neural Thermodynamic Laws for Large Language Model Training Statistical mechanics of learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.841719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.285139Z digest=sha256:eb436a4459671f932714ecfcd3b3a35e629644d063fb30d6e0c7c71d8a2cea6c

Observation 4b77da36-8744-449b-9ffa-0d35359d4f31 · outbound

This paper cites Statistical mechanics of deep learning.

Neural Thermodynamic Laws for Large Language Model Training Statistical mechanics of deep learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.829349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.288960Z digest=sha256:759d551997df6e078e691c2be55c47d25b6b7e3bea43409eb957dbd8f9ad16cd

Observation c0f021a9-d8f0-4eab-8225-95664e1e90bc · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Neural Thermodynamic Laws for Large Language Model Training Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.293933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.293933Z digest=sha256:0d1606ccc34ed2e96e40d86a875ca67c87cfb5edb815847c4277492deb17307c

Observation 11e2a7b6-6bce-48d9-8fae-5acc1ef8208e · outbound

This paper cites How noise affects the Hessian spectrum in overparameterized neural networks.

Neural Thermodynamic Laws for Large Language Model Training How noise affects the Hessian spectrum in overparameterized neural networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.299041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.299041Z digest=sha256:3892bc871e4db4fa0c8ac3f641725e5c76a211488bce6f0a3e053ef9f932dc11

Observation 5705197c-17c9-408b-b4a5-fb1d4f50b76b · outbound

This paper cites FOCUS: First Order Concentrated Updating Scheme.

Neural Thermodynamic Laws for Large Language Model Training FOCUS: First Order Concentrated Updating Scheme

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.303679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.303679Z digest=sha256:50e229f806e4c8a2b93f589d54789c8e3c278c552153ac89f702fbf5bf94c61f

Observation 3c1eea92-38df-4956-89e6-2fc88a336a71 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Neural Thermodynamic Laws for Large Language Model Training MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.308075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.308075Z digest=sha256:2638033b6869a0db9bc02c01b2ec7f7b5d3aa47bd131501e289072fb8d8f5291

Observation be573704-5082-4bfe-8460-94297b5b5030 · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Neural Thermodynamic Laws for Large Language Model Training Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.312287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.312287Z digest=sha256:47ed2c2310b15dfeb92aa6bbb6a9d650596418d83006de250e06dff348c245c7

Observation a58f5b92-0d8c-46dd-a841-f45d69571120 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.

Neural Thermodynamic Laws for Large Language Model Training Scaling laws and compute-optimal training beyond fixed training durations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.817150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.316507Z digest=sha256:a84b6c148631ffff3b16c453b229b7ae88af47a0405374c3a5062dc881799375

Observation 54d121ff-1a49-4fc7-bda8-ba7003fa8064 · outbound

This paper cites an unresolved cited work.

Neural Thermodynamic Laws for Large Language Model Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.320808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.320808Z digest=sha256:675115f5d073907d19047d85f6c886a9cd5eadfd900d055d08eec0f5d9056aea

Observation 5b9dbac8-217c-4ea7-8ed3-ea12abddc571 · outbound

This paper cites A multi-power law for loss curve prediction across learning rate schedules.

Neural Thermodynamic Laws for Large Language Model Training A multi-power law for loss curve prediction across learning rate schedules

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.795287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.324525Z digest=sha256:229e3ffad16ef114d430efe2cfb18a8a89298c485304cbb994aa303b18765bf3

Observation adb2480a-b973-45ce-8c44-4d3486f0c5e2 · outbound

This paper cites modded-nanogpt.

Neural Thermodynamic Laws for Large Language Model Training modded-nanogpt

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.781648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.328811Z digest=sha256:89a33404b47cd3a8d7a38dc92878d03a95511ab37b60ecb5d9bfe8c9671191e8

Observation e10bec04-7f0f-4a15-9eab-f4feff0a122d · outbound

This paper cites Implicit Gradient Regularization.

Neural Thermodynamic Laws for Large Language Model Training Implicit Gradient Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.333148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.333148Z digest=sha256:e84108718eea307925c81658fd6833ed65c6acd93569e3344751cb2c2b649121

Observation 975e5bbb-1674-458d-926f-1134ba17a91b · outbound

This paper cites The limiting dynamics of sgd: Modified loss, phase-space oscillations, and anomalous diffusion.

Neural Thermodynamic Laws for Large Language Model Training The limiting dynamics of sgd: Modified loss, phase-space oscillations, and anomalous diffusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.770762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.337208Z digest=sha256:18591ed6eb5cb71a74a1f1d7690735ad83d033d1b61f4c2724f3248356e6602c

Observation 3fa17f3c-0219-4868-9251-f6ded64f650d · outbound

This paper cites Stochastic collapse: How gra- dient noise attracts sgd dynamics towards simpler subnetworks.Advances in Neural Information Processing Systems, 36:35027–35063, 2023.

Neural Thermodynamic Laws for Large Language Model Training Stochastic collapse: How gra- dient noise attracts sgd dynamics towards simpler subnetworks.Advances in Neural Information Processing Systems, 36:35027–35063, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.757806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.342128Z digest=sha256:fe459af4de3460119d1cfa5315ed8f940cf9c2b5a5e80119900b870c5c5664d4

Observation 710f4bf1-39bc-443b-93ce-34b437911117 · outbound

This paper cites Stochastic gradient descent as approximate bayesian inference.

Neural Thermodynamic Laws for Large Language Model Training Stochastic gradient descent as approximate bayesian inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.744777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.346839Z digest=sha256:9a3b9e75fff1c047c3e96d990e6af5eccbb06489f62a3ad6cc82ed0e4e615f90

Observation 3e756cd0-1ae7-49aa-b88d-f3b250cec91c · outbound

This paper cites A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.

Neural Thermodynamic Laws for Large Language Model Training A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.351351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.351351Z digest=sha256:b897200c39f38890266635317c44d354b1dce79ddf6288145e4136f0a28056fd

Observation d89849d2-59be-42e7-80ce-2cfc42587224 · outbound

This paper cites Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate.

Neural Thermodynamic Laws for Large Language Model Training Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:16:10.538298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.356219Z digest=sha256:5e19fa0e428499fd8f7a2d056b3502ded37fe10386298f66b8d3fe19034157f6

Observation dbec16f3-8046-48dd-9e77-15f4b5af7ac0 · outbound

This paper cites Gradient Descent Maximizes the Margin of Homogeneous Neural Networks.

Neural Thermodynamic Laws for Large Language Model Training Gradient Descent Maximizes the Margin of Homogeneous Neural Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.360638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.360638Z digest=sha256:193c96975e69950f8df98a32d60ccddda1094c1be44e1fa656ff656fabbb9d72

Observation 1d8dc6a4-c009-475a-a7e6-526b7fd195d4 · outbound

This paper cites The implicit bias for adaptive optimization algorithms on homogeneous neural networks.

Neural Thermodynamic Laws for Large Language Model Training The implicit bias for adaptive optimization algorithms on homogeneous neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.732032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.364472Z digest=sha256:8ba57599ed2d5fa8e62d4073960d88154c8087cfed43c3f52db01fe8f3796018

Observation 5bf4fd29-7f62-4325-bc6a-6971e63bb0e1 · outbound

This paper cites An overview of condensation phenomenon in deep learning.

Neural Thermodynamic Laws for Large Language Model Training An overview of condensation phenomenon in deep learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.369421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.369421Z digest=sha256:bc3f7af466e3e0001f90caa24f7ee9a06cb9dd96d0c69435839e99769942b5b8

Observation d30cc1f6-f2d3-4bec-8fef-f587ab01b983 · outbound

This paper cites Loss surfaces, mode connectivity, and fast ensembling of dnns.

Neural Thermodynamic Laws for Large Language Model Training Loss surfaces, mode connectivity, and fast ensembling of dnns

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.373316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.373316Z digest=sha256:05461c25c425300bdb6352d645a07619bd8b374096d8fefddaad686ffee36fb2

Observation e6fee98e-5454-4dfd-98ee-63807f5f9ac9 · outbound

This paper cites Linear mode connectivity and the lottery ticket hypothesis.

Neural Thermodynamic Laws for Large Language Model Training Linear mode connectivity and the lottery ticket hypothesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.378385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.378385Z digest=sha256:7556fd29a7b6a039d532c00bc3bf0a5ed2c95790523e6f8505e1b19a225a8d83

Observation 14003317-7996-4b1c-aa40-4bcd29f84b95 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Neural Thermodynamic Laws for Large Language Model Training SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.382408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.382408Z digest=sha256:5eb882a2fcccc64f310b3c038a0881c6052d1e6827a369c93c31e5ecc6e62340

Observation c9c19445-75c5-4c60-b43b-ed7169e8d14d · outbound

This paper cites Cyclical learning rates for training neural networks.

Neural Thermodynamic Laws for Large Language Model Training Cyclical learning rates for training neural networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.703872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.386589Z digest=sha256:32837995e10df42ab662b94d0a0c97df4a10c18eb2617505637ab9875a8a34ed

Observation cae9e569-d9fb-44df-8f14-d605328e46c4 · outbound

This paper cites Attention is all you need.

Neural Thermodynamic Laws for Large Language Model Training Attention is all you need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.391341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.391341Z digest=sha256:2b2b3cbdbd858ca152cc902944697c14fc6a217742e412e9f9387f9b2a4548f8

Observation 389cbea2-8c0b-46ca-a545-57a66aebce39 · outbound

This paper cites The information bottleneck method.

Neural Thermodynamic Laws for Large Language Model Training The information bottleneck method

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.395835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.395835Z digest=sha256:0877d441689838a313b8a0d5bc9c3c178718ee810f195554d1450d39ebabce4e

Observation cf45d999-1565-418a-abed-41956b681423 · outbound

This paper cites Entropy-sgd: Biasing gradient descent into wide valleys.

Neural Thermodynamic Laws for Large Language Model Training Entropy-sgd: Biasing gradient descent into wide valleys

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.681046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.400168Z digest=sha256:1410eb1f0c42d42ed91cd43f7b170f6bbcef9860cad096ecfce6dbad67f7aeb9

Observation 0221e432-621f-4859-9a46-591124545197 · outbound

This paper cites A learning algorithm for boltzmann machines.

Neural Thermodynamic Laws for Large Language Model Training A learning algorithm for boltzmann machines

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.667575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.405438Z digest=sha256:adf5728ba03e0d845ae2774599fc2b4b9134a5f97ef00ea58331608c1fc3313f

Observation 98bdeea6-d6e2-4d6d-96de-4eb885ceffae · outbound

This paper cites Hopfield Networks is All You Need.

Neural Thermodynamic Laws for Large Language Model Training Hopfield Networks is All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.410063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.410063Z digest=sha256:c3334aad588810b96e00f98884797a7b038652d42fa422418cdf06a97b4ded9d

Observation 9c2a8453-6e38-4eb4-804f-834e09498e3a · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Neural Thermodynamic Laws for Large Language Model Training Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.414092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.414092Z digest=sha256:95793957858d709745ab41c58d20589ecd2248abc15f91fbd4517ed128e1020f

Observation 30cca370-659f-4db5-9d77-e0379a008b63 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Neural Thermodynamic Laws for Large Language Model Training Score-Based Generative Modeling through Stochastic Differential Equations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.417979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.417979Z digest=sha256:dd1b0d6971c619a6813c87f210c02d46155906a2c5a4c3f18d28ae93fb3ef986

Observation 8430a16a-5d30-48c7-a45d-43a0ab8ab73c · outbound

This paper cites edge of stability.

Neural Thermodynamic Laws for Large Language Model Training edge of stability

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:16:10.647787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:16:10.422836Z digest=sha256:6453b2fdeb53e6d4322a081d387b6cf4467f72fbde877e2d27cfe66ae2e03c52

Pith citing papers

Observation 52d8cbf4-a455-4428-b670-c19fcb8309d5 · inbound

Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model cites this paper.

Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model Neural Thermodynamic Laws for Large Language Model Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:05.518604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:05.518604Z digest=sha256:7b65a395fca9d39d9ecd967ccf3142e17fdd0b84c810d1e8b448d28bcfa32e27

Observation b88f9cc7-7416-4d8d-b8e8-91e37a2868d3 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Neural Thermodynamic Laws for Large Language Model Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:05.058346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:05.058346Z digest=sha256:ffba5bd55e99358a2926ece5cf1cfa6a3d40cdf20151ef7215a031fa806535c1

Observation 4617356d-24ea-480a-be5f-e57e497bff5e · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Neural Thermodynamic Laws for Large Language Model Training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:43.275595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:cf191077b2c8bfe47a5960b4553ae63bca7a38f28fe915e23cda10e8dd5d7898