Pith. sign in

Paper Citation Record · LEDGER

Neural Thermodynamic Laws for Large Language Model Training

As of 15 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 3 inbound Pith citation observations for arXiv:2505.10559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10559 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:16:10.422836Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:59:05.518604Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T07:00:43.273983Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d212e579-4736-4103-b4d3-8542c19ba52b · outbound

This paper cites Bayesian learning via stochastic gradient langevin dynamics.

Neural Thermodynamic Laws for Large Language Model Training Bayesian learning via stochastic gradient langevin dynamics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.276576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.276576Z digest=sha256:41e252aaf64f09e705a3b4cc3c79edbb80bee2b6862b4673ee9007198a362cc5

Observation 6a950008-b592-48e9-b839-c497e6d402ee · outbound

This paper cites Thermodynamics-inspired explanations of artificial intelli- gence.

Neural Thermodynamic Laws for Large Language Model Training Thermodynamics-inspired explanations of artificial intelli- gence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.854755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.281206Z digest=sha256:22bc2f975a54a2f7b8e6f255c090f9dc1e01dc976370978de7c61ac5a63c84d1

Observation eb823f02-bcc2-435c-b0db-c5cca54d8faa · outbound

This paper cites Statistical mechanics of learning.

Neural Thermodynamic Laws for Large Language Model Training Statistical mechanics of learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.841719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.285139Z digest=sha256:64ed352475c1c0da9245f554496caff9bb1ff30dc04c125b1dcbda0ffcc14dee

Observation 4b77da36-8744-449b-9ffa-0d35359d4f31 · outbound

This paper cites Statistical mechanics of deep learning.

Neural Thermodynamic Laws for Large Language Model Training Statistical mechanics of deep learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.829349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.288960Z digest=sha256:e965970de72ee6a677457af7acb90552b57107cda64cbdcf3d6b5492ca4515a3

Observation c0f021a9-d8f0-4eab-8225-95664e1e90bc · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Neural Thermodynamic Laws for Large Language Model Training Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.293933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.293933Z digest=sha256:29cc4caa12a5ed53b249e60eb150fb8ecc07ac088ecf12ff876ed95dc16f2894

Observation 11e2a7b6-6bce-48d9-8fae-5acc1ef8208e · outbound

This paper cites How noise affects the Hessian spectrum in overparameterized neural networks.

Neural Thermodynamic Laws for Large Language Model Training How noise affects the Hessian spectrum in overparameterized neural networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.299041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.299041Z digest=sha256:2b2460c28e4831e008b5cabdd84c134c489c0a291cfefd1ceedbb73df459ccb9

Observation 5705197c-17c9-408b-b4a5-fb1d4f50b76b · outbound

This paper cites FOCUS: First Order Concentrated Updating Scheme.

Neural Thermodynamic Laws for Large Language Model Training FOCUS: First Order Concentrated Updating Scheme

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.303679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.303679Z digest=sha256:4e18791d88403694ca2a37df301583df01869df93aea3b2979e19b9156f4294d

Observation 3c1eea92-38df-4956-89e6-2fc88a336a71 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Neural Thermodynamic Laws for Large Language Model Training MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.308075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.308075Z digest=sha256:e7795c5500d8e3840f690cdcc76b0878f404589bfde0eec39dc54260e6341668

Observation be573704-5082-4bfe-8460-94297b5b5030 · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Neural Thermodynamic Laws for Large Language Model Training Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.312287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.312287Z digest=sha256:bda5f24990b6481ca3b6b18fff21a3a0a94bd917e77f5e7b60cac3bca2ccfe9c

Observation a58f5b92-0d8c-46dd-a841-f45d69571120 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.

Neural Thermodynamic Laws for Large Language Model Training Scaling laws and compute-optimal training beyond fixed training durations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.817150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.316507Z digest=sha256:2b61bd90db8d03da7661c1e5a046268bdab2d5054139328685b157f1176924c6

Observation 54d121ff-1a49-4fc7-bda8-ba7003fa8064 · outbound

This paper cites an unresolved cited work.

Neural Thermodynamic Laws for Large Language Model Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.320808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.320808Z digest=sha256:5c7b1538821dafb0e5b2a4b6e3da6aa001493b9b90f0b03b1fbf454193250b9b

Observation 5b9dbac8-217c-4ea7-8ed3-ea12abddc571 · outbound

This paper cites A multi-power law for loss curve prediction across learning rate schedules.

Neural Thermodynamic Laws for Large Language Model Training A multi-power law for loss curve prediction across learning rate schedules

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.795287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.324525Z digest=sha256:b5886f4d496c664bc9e0243e292bfcda205dc29e1a4fc8c2b4508897cc8302dc

Observation adb2480a-b973-45ce-8c44-4d3486f0c5e2 · outbound

This paper cites modded-nanogpt.

Neural Thermodynamic Laws for Large Language Model Training modded-nanogpt

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.781648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.328811Z digest=sha256:85ac37f70c58cf1511c551be14ab729dcd15cffd5c6bc4b4f0d1397f771dd829

Observation e10bec04-7f0f-4a15-9eab-f4feff0a122d · outbound

This paper cites Implicit Gradient Regularization.

Neural Thermodynamic Laws for Large Language Model Training Implicit Gradient Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.333148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.333148Z digest=sha256:46049cc86f248e9525e7463d5239e894ebab20343c8fbaee5d082483b15db7f6

Observation 975e5bbb-1674-458d-926f-1134ba17a91b · outbound

This paper cites The limiting dynamics of sgd: Modified loss, phase-space oscillations, and anomalous diffusion.

Neural Thermodynamic Laws for Large Language Model Training The limiting dynamics of sgd: Modified loss, phase-space oscillations, and anomalous diffusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.770762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.337208Z digest=sha256:63503f422e8aeaa58c6c79175ace2eb1f64d59f2ef0e3973281962a9570d6263

Observation 3fa17f3c-0219-4868-9251-f6ded64f650d · outbound

This paper cites Stochastic collapse: How gra- dient noise attracts sgd dynamics towards simpler subnetworks.Advances in Neural Information Processing Systems, 36:35027–35063, 2023.

Neural Thermodynamic Laws for Large Language Model Training Stochastic collapse: How gra- dient noise attracts sgd dynamics towards simpler subnetworks.Advances in Neural Information Processing Systems, 36:35027–35063, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.757806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.342128Z digest=sha256:8285d8c21b52f6b8f370e7fa8ba69d9c83864e98e916c98d39dc8f0f81887bc6

Observation 710f4bf1-39bc-443b-93ce-34b437911117 · outbound

This paper cites Stochastic gradient descent as approximate bayesian inference.

Neural Thermodynamic Laws for Large Language Model Training Stochastic gradient descent as approximate bayesian inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.744777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.346839Z digest=sha256:84682bad29ed720001d1ca739f980b59df618faa55e3c406aad0b17aca2e48a2

Observation 3e756cd0-1ae7-49aa-b88d-f3b250cec91c · outbound

This paper cites A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.

Neural Thermodynamic Laws for Large Language Model Training A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.351351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.351351Z digest=sha256:c82547b16a511503a7fe3b1ef8cbb755c26cc8b076c49b91d64a9df7f8bc8cbb

Observation d89849d2-59be-42e7-80ce-2cfc42587224 · outbound

This paper cites Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate.

Neural Thermodynamic Laws for Large Language Model Training Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:16:10.538298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.356219Z digest=sha256:db3a870723560150269fbcdfe53b12efb5783accc25c8f9bb3e018debf7a5b6d

Observation dbec16f3-8046-48dd-9e77-15f4b5af7ac0 · outbound

This paper cites Gradient Descent Maximizes the Margin of Homogeneous Neural Networks.

Neural Thermodynamic Laws for Large Language Model Training Gradient Descent Maximizes the Margin of Homogeneous Neural Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.360638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.360638Z digest=sha256:3796f54af4de3ba957fee259687ae181d4452fa99aaa428596c93acb2b3e387c

Observation 1d8dc6a4-c009-475a-a7e6-526b7fd195d4 · outbound

This paper cites The implicit bias for adaptive optimization algorithms on homogeneous neural networks.

Neural Thermodynamic Laws for Large Language Model Training The implicit bias for adaptive optimization algorithms on homogeneous neural networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.732032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.364472Z digest=sha256:0d2739f85a3e6de3a31320f1d5c6c10b44f832df7d353503db4fe626bff6573a

Observation 5bf4fd29-7f62-4325-bc6a-6971e63bb0e1 · outbound

This paper cites An overview of condensation phenomenon in deep learning.

Neural Thermodynamic Laws for Large Language Model Training An overview of condensation phenomenon in deep learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.369421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.369421Z digest=sha256:12cc403946291214ea474b75bcd7adbd1ca07b516c738e282b96f2090bc19057

Observation d30cc1f6-f2d3-4bec-8fef-f587ab01b983 · outbound

This paper cites Loss surfaces, mode connectivity, and fast ensembling of dnns.

Neural Thermodynamic Laws for Large Language Model Training Loss surfaces, mode connectivity, and fast ensembling of dnns

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.373316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.373316Z digest=sha256:60f97f19d9c0fef69f238a6c3e2a19b5e90110382caa791064807b75410905a7

Observation e6fee98e-5454-4dfd-98ee-63807f5f9ac9 · outbound

This paper cites Linear mode connectivity and the lottery ticket hypothesis.

Neural Thermodynamic Laws for Large Language Model Training Linear mode connectivity and the lottery ticket hypothesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.378385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.378385Z digest=sha256:f51da8881e819322087db837d7fc8f417ca18be17404a1050e69d36b1107c711

Observation 14003317-7996-4b1c-aa40-4bcd29f84b95 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Neural Thermodynamic Laws for Large Language Model Training SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.382408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.382408Z digest=sha256:234f750fa18b7fe05549901010956d61f51b5293fefe9e2b9f9dd6ec769ddebf

Observation c9c19445-75c5-4c60-b43b-ed7169e8d14d · outbound

This paper cites Cyclical learning rates for training neural networks.

Neural Thermodynamic Laws for Large Language Model Training Cyclical learning rates for training neural networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.703872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.386589Z digest=sha256:f36e138ea03cb2516ca46ca0aea23cfa5a8d03af2453922b899ccc38335b85d9

Observation cae9e569-d9fb-44df-8f14-d605328e46c4 · outbound

This paper cites Attention is all you need.

Neural Thermodynamic Laws for Large Language Model Training Attention is all you need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.391341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.391341Z digest=sha256:8a14f6c99030323ac0875a2b40b31011d5cca89ea8d9990589393f1fb3ddce30

Observation 389cbea2-8c0b-46ca-a545-57a66aebce39 · outbound

This paper cites The information bottleneck method.

Neural Thermodynamic Laws for Large Language Model Training The information bottleneck method

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.395835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.395835Z digest=sha256:20a99fa8c6ac5e31d3701fd61fb3a0bc292fd25a15a12c2a75cdda1aaf186da3

Observation cf45d999-1565-418a-abed-41956b681423 · outbound

This paper cites Entropy-sgd: Biasing gradient descent into wide valleys.

Neural Thermodynamic Laws for Large Language Model Training Entropy-sgd: Biasing gradient descent into wide valleys

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.681046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.400168Z digest=sha256:89f4b66b3083f07bfba4cd4d3c958b38086bf7dbf8266800183c79752b57725c

Observation 0221e432-621f-4859-9a46-591124545197 · outbound

This paper cites A learning algorithm for boltzmann machines.

Neural Thermodynamic Laws for Large Language Model Training A learning algorithm for boltzmann machines

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:16:10.667575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.405438Z digest=sha256:edcb50621525c77e8e0caefa0a75995f8b177745b3d627130304c7b7808e079a

Observation 98bdeea6-d6e2-4d6d-96de-4eb885ceffae · outbound

This paper cites Hopfield Networks is All You Need.

Neural Thermodynamic Laws for Large Language Model Training Hopfield Networks is All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.410063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.410063Z digest=sha256:d526e94875d58dd040dd18d6a389fc78f40f974d105ada934297a4243b048d70

Observation 9c2a8453-6e38-4eb4-804f-834e09498e3a · outbound

This paper cites Deep unsuper- vised learning using nonequilibrium thermodynamics.

Neural Thermodynamic Laws for Large Language Model Training Deep unsuper- vised learning using nonequilibrium thermodynamics

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.414092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.414092Z digest=sha256:22e59893d7a6285cc8c3d89a28c73329290a07e15e2e265a6c6286a25cb73886

Observation 30cca370-659f-4db5-9d77-e0379a008b63 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Neural Thermodynamic Laws for Large Language Model Training Score-Based Generative Modeling through Stochastic Differential Equations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:16:10.417979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:16:10.417979Z digest=sha256:701be5c9b207860e6245c6ff69e7d54505c20c874500abe3be584364a0763e6f

Observation 8430a16a-5d30-48c7-a45d-43a0ab8ab73c · outbound

This paper cites edge of stability.

Neural Thermodynamic Laws for Large Language Model Training edge of stability

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:16:10.647787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T21:16:10.422836Z digest=sha256:5dae8f082e4fc5000040fa9aa1eaecb5d167cf28f56af23be9b597887c7f6de0

Pith citing papers

Observation 52d8cbf4-a455-4428-b670-c19fcb8309d5 · inbound

Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model cites this paper.

Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model Neural Thermodynamic Laws for Large Language Model Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:05.518604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:05.518604Z digest=sha256:66e0a84e45d96d9b43f6404557c4836586de1f26c8b7ca4f894a6d1f3499b1ab

Observation b88f9cc7-7416-4d8d-b8e8-91e37a2868d3 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Neural Thermodynamic Laws for Large Language Model Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:05.058346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:05.058346Z digest=sha256:465799d0d291078ff2e7db29b47dd10568c12dbd575296c39cbe3f685af387ec

Observation 4617356d-24ea-480a-be5f-e57e497bff5e · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Neural Thermodynamic Laws for Large Language Model Training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:43.275595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:97881d8f2d199ac6aee772e6fc479484af37757cf24ba29ffd7da8da6bb00e70