Pith. sign in

Paper Citation Record · LEDGER

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training

As of 14 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2505.23489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23489 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:46.143996Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact8
  • verified fuzzy42
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c5405bf-6824-4880-a578-313cde9c8185 · outbound

This paper cites TherML: Thermodynamics of Machine Learning.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training TherML: Thermodynamics of Machine Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:48:50.002829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:36.572374Z digest=sha256:665bed58d88e650a37d1587f7af72060455d926e8fb6174380ffb2478ea80e1f

Observation 3243b1bb-8435-4704-81a1-041a34cd44dd · outbound

This paper cites SGD with Large Step Sizes Learns Sparse Features.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training SGD with Large Step Sizes Learns Sparse Features

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:49.648252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:36.710873Z digest=sha256:44707e10d01f41c0eaeaf1b17ab8d0448cc7c82577c3deb2b044920820b01aa7

Observation 517aaa9f-c710-4dc6-b5a0-076d38697fad · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:36.852990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:36.852990Z digest=sha256:b2bc2ec41b3d43909d97e9361d143d9581197c917c903aea584e006941e932ce

Observation 2fdbb3af-20d8-42ed-897c-9cb699d81e6e · outbound

This paper cites Implicit gradient regularization.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Implicit gradient regularization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:02.405254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:36.995113Z digest=sha256:1db3f76eab04523a39c73a49003e227f887751a079c5c53eef3dbd5014d132ff

Observation 6b82d6ee-f178-472c-ba2f-663d27710d13 · outbound

This paper cites Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Science, 116(32):15849–15854, 2019.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Science, 116(32):15849–15854, 2019

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:49:02.156661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:37.146064Z digest=sha256:85cd73cc243df55572f3feef8f65d1bacfb1560260e358f351665a745f21ce22

Observation 288f3599-7164-4196-a634-048bf1de99ce · outbound

This paper cites Practical recommendations for gradient-based training of deep architectures.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Practical recommendations for gradient-based training of deep architectures

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:37.300545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:37.300545Z digest=sha256:b5b70705e0a14eb50bb05b978492ca204728a6f1a3edd950213bb1256cc6bf2b

Observation fefda19d-bf62-476f-9342-a81d024e532b · outbound

This paper cites Language Models are Few-Shot Learners.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:37.430097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:37.430097Z digest=sha256:6e8d3b8a17e154ad1561f9239d9c5227524aa8b3595fd9134b5db5ded51debee

Observation e3849048-9ae6-4931-9f0a-257e1a7a22f1 · outbound

This paper cites Stochastic gradient descent performs variational infer- ence, converges to limit cycles for deep networks.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic gradient descent performs variational infer- ence, converges to limit cycles for deep networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:01.941909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:37.613521Z digest=sha256:c1bcad555c427f7770c4f1f0d5b17eceb0800e311abe93ba8de6c9d62b32f651

Observation fa6adc37-e9a1-4333-81b1-036b7b5ec6e9 · outbound

This paper cites Entropy-SGD: Biasing gradient descent into wide valleys.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Entropy-SGD: Biasing gradient descent into wide valleys

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:01.606523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:37.787127Z digest=sha256:4758ab14dd00a0feb5d72ccab7f3ca057af1a65b18cec369e1c0d8448f620d0c

Observation e3b06684-69c2-45c8-8162-1cc2a6d9f69b · outbound

This paper cites Convergence diagnostics for stochastic gradient descent with constant learning rate.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Convergence diagnostics for stochastic gradient descent with constant learning rate

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:01.326095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:37.938749Z digest=sha256:54b18d4c6468ec5ce3544923b7dae7385eafd68743002187028ad6f20049d478

Observation 961ad5b2-8ec7-4673-bfc3-d1623dc07ab7 · outbound

This paper cites Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:01.066149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.046351Z digest=sha256:9abc7982317019f26a49eee6ed30c0aaf400dee9cc8db4eb2a144dc95d93e47b

Observation cc138851-290c-4b68-8622-d741abb879fc · outbound

This paper cites Stochastic collapse: How gra- dient noise attracts SGD dynamics towards simpler subnetworks.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic collapse: How gra- dient noise attracts SGD dynamics towards simpler subnetworks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:00.734116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.125374Z digest=sha256:52f6f3226a66f72377336baf1e8f8fde7dc7507b8f91aecb3f2fdb2049ab68e9

Observation 7422fef2-f593-4b54-942f-20077021b893 · outbound

This paper cites Symbolic discovery of optimization algorithms.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Symbolic discovery of optimization algorithms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:00.460505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.200072Z digest=sha256:8da5cf30c32237ca5055627f104dec0a80b9e9d30cabb923e6a5a6e21a6367f4

Observation 16d05eca-ae36-4b1d-8361-dbe6d71bc8d8 · outbound

This paper cites Gradient descent on neural networks typically occurs at the edge of stability.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Gradient descent on neural networks typically occurs at the edge of stability

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:00.159172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.281967Z digest=sha256:7c282b6c59e88692c0d936ef02776a2f0a177d10c076155e44a24e43e1c63831

Observation 2b1a750f-4c4b-471b-ad38-a82be506ce70 · outbound

This paper cites URL https://constructor.tech/products/ research-platform.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training URL https://constructor.tech/products/ research-platform

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:59.887943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.353729Z digest=sha256:9af39801b068793a5a2e01683e6af51c0b69b0eed0bcd2075355f8203883c8d4

Observation 966fca58-2736-4112-81f1-87aba74e8ed3 · outbound

This paper cites Determining intrinsic dimension and entropy of high-dimensional shape spaces.Modeling and Simulation in Science, Engineering and Technology, pages 231–252,.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Determining intrinsic dimension and entropy of high-dimensional shape spaces.Modeling and Simulation in Science, Engineering and Technology, pages 231–252,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:59.602661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.448719Z digest=sha256:116610af066e618287cdaf245fbccd43f9f5643edd081dd8d91682c4df0337bc

Observation 1a6b142d-03bf-4d07-b6f5-f8c36e64625b · outbound

This paper cites Why do we need weight decay in modern deep learning? InAdvances in Neural Information Processing Systems, 2024.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Why do we need weight decay in modern deep learning? InAdvances in Neural Information Processing Systems, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:59.338853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.626357Z digest=sha256:d49e8a776da611e06f1f3fe27d07e3a4a8384e2eb0ea2edd696f1f5e91e94b40

Observation 91df84ab-a8df-416c-aae8-7be012c5531e · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:59.129696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.720001Z digest=sha256:6098cbf6a375d66506edf15cc43e896c51baecf4196d09ce920ac4a7ff68c9c9

Observation 0918aa1a-5731-47e4-8aa1-ed4a0a3bc1fa · outbound

This paper cites Essentially no barriers in neural network energy landscape.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Essentially no barriers in neural network energy landscape

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:58.803625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.899024Z digest=sha256:d6401ab52c16f0f5ad0e00e96169cdf338752fb62f2a66cd7e8ce6f40235291f

Observation a19d1065-d8b1-46db-8f53-43c663845821 · outbound

This paper cites Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:58.652430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:39.121524Z digest=sha256:6eee62237dabb3e98d05be39df7779ea32861263a443a16f5c4b1aed77516370

Observation 76675590-5794-4a6a-8824-09adfc9074bd · outbound

This paper cites A free-energy principle for representation learning.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training A free-energy principle for representation learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:58.431590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:39.323634Z digest=sha256:32e47a9d3491f944ccc2793b9ddba9db0db27ad4c18d4b7b0bdc9503831da0cd

Observation d8a40732-1810-4aba-8157-d1c6b296f298 · outbound

This paper cites Fixed-time stable gradient flows: Applications to continuous- time optimization.IEEE Transactions on Automatic Control, 66(5):2002–2015, 2021.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Fixed-time stable gradient flows: Applications to continuous- time optimization.IEEE Transactions on Automatic Control, 66(5):2002–2015, 2021

Reference 22

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:48:49.309456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:39.510629Z digest=sha256:cf9e6facf17ef545396d023399021346d1697694ae08969922d0a796a3d80e99

Observation 8b483a68-5559-4e66-8690-821feaafcd8d · outbound

This paper cites Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:39.612467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:39.612467Z digest=sha256:e7fce254940f1614d14d16af33d641116b66ef8ad969adab087af4ea5929253c

Observation 576d4845-946c-43f6-86f6-1da01bcffbb9 · outbound

This paper cites Stochastic training is not necessary for generalization.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic training is not necessary for generalization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:58.252193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:39.777083Z digest=sha256:efeea688acd8088746bd0502a0b761496d615e114c1254e09e25049fc1a3bf10

Observation e52ca576-43c7-40b2-81f5-4ab1eecbe448 · outbound

This paper cites Abrupt learning in transformers: A case study on matrix completion.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Abrupt learning in transformers: A case study on matrix completion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:57.896204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:39.902009Z digest=sha256:c565b41deae61b720a4f39577087b3e70450e3b63782eb60bb3f928ddd1b70bc

Observation 93369438-8a66-4d8b-96bc-4e3ecda53feb · outbound

This paper cites Deep Residual Learning for Image Recognition.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep Residual Learning for Image Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:39.993026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:39.993026Z digest=sha256:c530171630eaed3a25a470a32a54e1d088fdc4e6e6abe0dc3933a1fd8d7620f7

Observation 73431ffa-d71a-4c42-8e26-5b6a9fba3ed9 · outbound

This paper cites Three Factors Influencing Minima in SGD.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Three Factors Influencing Minima in SGD

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:40.106339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:40.106339Z digest=sha256:aa7fb6c1aba3eaddb0350f3569c55122a4594d4370309725def3ca6ba2ff750b

Observation 25c204db-ca40-4720-a820-5da35deaa3fd · outbound

This paper cites Scaling Laws for Neural Language Models.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Scaling Laws for Neural Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:40.245350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:40.245350Z digest=sha256:a5b2e14366b6831942abae7e32d794c0e17fc28bd9a6dbd65dde38ada57db993

Observation 2015c824-c544-4b57-b3cf-11d281e975b5 · outbound

This paper cites Kingma and Jimmy Ba.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Kingma and Jimmy Ba

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:57.653712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:40.382897Z digest=sha256:61662b76d77dce0b179ab7e32bcfb359695e90675f16357f09b76eea106674a1

Observation 231877f2-6a5b-4758-aa3a-d070439aed62 · outbound

This paper cites Training scale-invariant neural networks on the sphere can happen in three regimes.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Training scale-invariant neural networks on the sphere can happen in three regimes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:57.320267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:40.449755Z digest=sha256:8698a039568327d7db845c1f134099d624204ad907c5a5defd6f4443e037d628

Observation 4d1962e9-77e7-4320-8ddd-65a345c5c553 · outbound

This paper cites Big Transfer (BiT): General Visual Representation Learning.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Big Transfer (BiT): General Visual Representation Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:40.523882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:40.523882Z digest=sha256:66683668a929875a91e92afca2966205e046b372c73edb18316d8cdbc1b38f4b

Observation eee4deaa-30ca-43ae-a0b2-0d42a487b4d4 · outbound

This paper cites CIFAR-10 (canadian institute for advanced research).

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training CIFAR-10 (canadian institute for advanced research)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:57.030510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:40.657045Z digest=sha256:4c8e4f3ab5380bd674eda5ba1134f2f2c984466e63fe993a265c502615a63cb0

Observation 32b601c5-e9d5-4cd7-9a81-e58de55d9de6 · outbound

This paper cites CIFAR-100 (canadian institute for advanced research).

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training CIFAR-100 (canadian institute for advanced research)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.712743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:40.746665Z digest=sha256:270ffc1f6efc4a8650dcb79c71dfea78be81c3add5facd6b5a3b3369c0a9fbe3

Observation 19d3eb95-9c45-4b8a-be47-56a9d7b1d7f9 · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T12:48:47.794906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:40.831671Z digest=sha256:3e3187d4cd5f6da72f4dc628469094c6f86a3732124913fd18ed8d47db779526

Observation e8891303-7c11-46e5-a72a-ea91e4b653c2 · outbound

This paper cites Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:40.958072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:40.958072Z digest=sha256:b413fce7173616c3c91d3c2074974b0951b88300726a5df09845f609d092a06b

Observation ca1d1278-3582-4bfe-94dc-f0c1cdf38cfe · outbound

This paper cites Few-shot adaptation of multi-modal foundation models: A survey.Artificial Intelli- gence Review, 57(10):268, 2024.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Few-shot adaptation of multi-modal foundation models: A survey.Artificial Intelli- gence Review, 57(10):268, 2024

Reference 36

Resolution
verified exact
doi, observed 2026-08-07T12:48:47.484013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:41.143952Z digest=sha256:ba66f381af8595b8de512257522dc4732c5807fc08781d6fb252326cc4dfdf4a

Observation f575c99c-1242-4d96-af71-e05728df534e · outbound

This paper cites Understanding why neural networks generalize well through GSNR of parameters.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Understanding why neural networks generalize well through GSNR of parameters

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.450939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:41.225917Z digest=sha256:face6685d65b9e9f2731177b42a80fbd72291d279b21282e324f6d4a4ad28989

Observation 8241ca34-06a4-40c1-914d-6141f0766561 · outbound

This paper cites Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.941810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:41.344996Z digest=sha256:a314b3ba08899ef503b81ab6624accb3b6a9aa8c32a6c965c68191b6ceddb351

Observation 7c6e4e48-6a9e-4715-b430-90b2723e4a71 · outbound

This paper cites Towards understanding grokking: An effective theory of representation learning.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Towards understanding grokking: An effective theory of representation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.234290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:41.465642Z digest=sha256:da0eb10c54c12e52468b7f0efc7958d4e64bb099c24085f235387702fdc0f65c

Observation 827a797e-dcfe-4010-befe-06364c2ae4b7 · outbound

This paper cites On the periodic behavior of neural network training with batch normalization and weight decay.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the periodic behavior of neural network training with batch normalization and weight decay

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.036199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:41.589945Z digest=sha256:50144def21a85f0b28bf6e915a32ead0a2f4e68c2cea8f5beb2b919660a9df33

Observation ebebddeb-4929-498a-92c9-45e220a39e4c · outbound

This paper cites Decoupled weight decay regularization.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Decoupled weight decay regularization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:41.701378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:41.701378Z digest=sha256:44476bca30333097e6434a84395dfd627c13fa264a8f71471fd00e2e87a72ee2

Observation 828f5cd0-057f-4141-a457-5f9059fdb4be · outbound

This paper cites The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:41.855492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:41.855492Z digest=sha256:652f4ab7e1b2a736584608290caeb342eb45e4c5db216eba3bf9ba9959633266

Observation 77d1edcc-38ab-401e-b0ef-e3fa9d5b9094 · outbound

This paper cites Stochastic Gradient Descent as Approximate Bayesian Inference.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic Gradient Descent as Approximate Bayesian Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:41.985527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:41.985527Z digest=sha256:6a3e7a419b4e78ecaf67e00d48fbf32b6ada028bd5752f1d65ffbbd1a54da238

Observation dbff6cb6-7569-4135-971f-733035b16bc0 · outbound

This paper cites Phase transitions in the mini-batch size for sparse and dense two-layer neural networks.Machine Learning: Science and Technology, 5(1): 015015, 2024.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Phase transitions in the mini-batch size for sparse and dense two-layer neural networks.Machine Learning: Science and Technology, 5(1): 015015, 2024

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:48:47.186730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:42.178037Z digest=sha256:ca5ba8244b3689f5cb3f6582e0d574dd98e04187832531a2d645b86de2758627

Observation e3d4a475-5726-4fab-a6a2-aa39e708545e · outbound

This paper cites Stochastic Gradient Descent on Separable Data: Exact Convergence with a Fixed Learning Rate.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic Gradient Descent on Separable Data: Exact Convergence with a Fixed Learning Rate

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:48.600944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:42.355392Z digest=sha256:54928c36c1aa2b10b360094ac0cad6671ea2267d42690bf97de946976eb8befd

Observation c781f22f-abda-448a-8f2d-7ad84fdf099d · outbound

This paper cites Bayesian Free Energy of Deep ReLU Neural Network in Overparametrized Cases.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Bayesian Free Energy of Deep ReLU Neural Network in Overparametrized Cases

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:42.512102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:42.512102Z digest=sha256:aabe93c79e6b1963ee0785975f26087962719eeedfa29c6b7f355aa6bacea481

Observation 56031d6d-5de4-4f21-af93-d6af4d27f522 · outbound

This paper cites Deep double descent: Where bigger models and more data hurt.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep double descent: Where bigger models and more data hurt

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.828535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:42.607428Z digest=sha256:507e682fdbb28676deb6063b12a0f5807d09fb739eeb3a9c124f13b3c892af15

Observation 7ffc897f-72ea-47a6-bea6-e32e919e7dbe · outbound

This paper cites LR0.FM: Low-resolution zero-shot classification benchmark for foundation models.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training LR0.FM: Low-resolution zero-shot classification benchmark for foundation models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.617982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:42.722384Z digest=sha256:c96942929dff2e72071a26cb15d9ccd57e8242e337214eb9e7c9b177dfc15637

Observation 1de94719-1f6a-4d71-bf51-888bc13cb8d8 · outbound

This paper cites Loss landscape: SGD has a better view.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Loss landscape: SGD has a better view

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.566771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:42.817525Z digest=sha256:2f0fee1cfa3142730a5002e0e387ac7f61e4ead21e81c8129d956a781aefeb5f

Observation d099425f-38a9-4378-810d-088fb2301e2c · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:42.923436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:42.923436Z digest=sha256:9e8d451eeacdec2b09769011a46392276820bfa891469e57485ec040722a164c

Observation 37f347c5-f46a-48a5-9ea0-de6afe4c3d50 · outbound

This paper cites Accelerating Large Batch Training via Gradient Signal to Noise Ratio (GSNR).

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Accelerating Large Batch Training via Gradient Signal to Noise Ratio (GSNR)

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:43.019137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:43.019137Z digest=sha256:ecd3e18af2c1a764e4979a59e7ba5b44f791f571fa91242e8fe376a400f91cd3

Observation cf4970ed-0469-497f-8a72-61ba9cb8f5d3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Learning transferable visual models from natural language supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.384360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.098578Z digest=sha256:ef3767c1763efd4c6a46d8246db8330872ea59527f5d426064b6d310af9914c9

Observation c01f840f-b890-48a8-9bd6-8d601ae36896 · outbound

This paper cites Where do large learning rates lead us? InAdvances in Neural Information Processing Systems, 2024.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Where do large learning rates lead us? InAdvances in Neural Information Processing Systems, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:55.111425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.211101Z digest=sha256:ae4977322a6b4c37662d595c394c9ce03127a4e9719d59d1a49dec993660d600

Observation 0e38b7c4-cdf5-48af-8f49-a64a26b5ae39 · outbound

This paper cites On the different regimes of stochastic gradient descent.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the different regimes of stochastic gradient descent

Reference 54

Resolution
verified exact
doi, observed 2026-08-07T12:48:46.818627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.295654Z digest=sha256:a6c52cf437ceb0cf18ffafc5257a1bad5a4ee97fd4bb8181d46625ba8271c35c

Observation 8bff367d-c3eb-4f2e-9e81-ef15edefd85c · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:54.855090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.496452Z digest=sha256:d0217612a5812ae5c9985e30d6ec686fd4a41d75a6cdac32bbbdf984821bf2d1

Observation 7d61f117-34a8-45cb-ae16-2c4881bc476d · outbound

This paper cites On the Generalization Benefit of Noise in Stochastic Gradient Descent.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the Generalization Benefit of Noise in Stochastic Gradient Descent

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:48:48.284766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.574615Z digest=sha256:60e56a951ec6643cee5a84b7f4ae9bcb35fe64ad35f10f6f95ddf43f8c2ef114

Observation f6fec504-538a-4709-8a10-96d443fc7f7d · outbound

This paper cites On the origin of implicit regular- ization in stochastic gradient descent.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the origin of implicit regular- ization in stochastic gradient descent

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.580089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.678452Z digest=sha256:4205259512152d717e858d79b06ba573a765628188ffdab4e0382223a0841622

Observation 53d40a8b-67e4-44b8-a71e-b3fa1970713e · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.Transactions on Machine Learning Research, 2023.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.Transactions on Machine Learning Research, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:54.032448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.941915Z digest=sha256:06ac70c8c4005a6729be631886b3ae296472a85aa958bc2b764f002eaf7e859a

Observation f553db86-34a5-4eb8-bc03-012018789d0a · outbound

This paper cites Unleashing the power of gradient signal-to-noise ratio for zero-shot NAS.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unleashing the power of gradient signal-to-noise ratio for zero-shot NAS

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:44.069104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:44.069104Z digest=sha256:3ece893e1a0723e7763f4dafd19e2ebbaf678aa5c0d992e7dea1246ad0ba58b6

Observation 328c658d-3707-4117-8e0a-39fdf8de6db0 · outbound

This paper cites Deep learning and the information bottleneck principle,.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep learning and the information bottleneck principle,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.747474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.188196Z digest=sha256:477bab4571fe2d44daf262cff9ec133286f6ae8286ca09479860960e06d5dc05

Observation c6382461-7c03-4953-8b2a-bf1fc8f2f61e · outbound

This paper cites The discovery of superconductivity.Physics Today, 63(9):38–43,.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training The discovery of superconductivity.Physics Today, 63(9):38–43,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.425800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.366536Z digest=sha256:afe97f25ab01ac1fa162d75bc16ec3ef5b3aa9d4f52521c6b4ea991d35331481

Observation c81ca0d2-a013-4afd-9237-66511ad32f92 · outbound

This paper cites A survey of basic thermodynamics, 2004.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training A survey of basic thermodynamics, 2004

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:53.156648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.586912Z digest=sha256:fad3e4700f764ff9b82a20c8f21861912c0714787b96169ed57032967b9c0784

Observation a41a7fe5-488a-4fe5-99fa-cfca3b4be752 · outbound

This paper cites Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.817300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.741482Z digest=sha256:f6ba13f22c9dc29c84c99e243ac802c23d00c894d9807463b75ae262acf82b49

Observation 6e7e1725-9700-4565-b8c1-8edda39592b0 · outbound

This paper cites Bayesian learning via stochastic gradient langevin dynamics.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Bayesian learning via stochastic gradient langevin dynamics

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.531469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.876386Z digest=sha256:58575985b815f6709ded7c507abef1d8e69f0c7d207efa16cdddf0dc4bb198b4

Observation d38b50ff-745b-4cc9-a359-5708d3fc7698 · outbound

This paper cites Towards few- shot adaptation of foundation models via multitask finetuning.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Towards few- shot adaptation of foundation models via multitask finetuning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:52.147253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.979783Z digest=sha256:e171d6eef38ab4f7086f16271718bfdcf3186de867e4bb646e6460894bc85203

Observation fc4b5a34-80a4-473f-af50-29d8dd8d6472 · outbound

This paper cites Fluctuation-dissipation relations for stochastic gradient descent.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Fluctuation-dissipation relations for stochastic gradient descent

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.794687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.148706Z digest=sha256:4a841ffbdf7c4a0f47500c3fc82b97a20e52a15c3307653aa6882d7244be0da3

Observation faf4f39e-dd6a-4079-a3d0-c9e1dd180e17 · outbound

This paper cites How Does Learning Rate Decay Help Modern Neural Networks?.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training How Does Learning Rate Decay Help Modern Neural Networks?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:45.281153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:45.281153Z digest=sha256:e841bb4ac5717599dee873e65de69e98e1cdd3a290c288b7235644b0963d0a3c

Observation 7b3fab21-76e1-4d7e-ba52-3c684d880759 · outbound

This paper cites Saxe, Madhu S.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Saxe, Madhu S

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:45.411715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:45.411715Z digest=sha256:d405ee4ac7a78fd2bf14cfc0066ddfc6b71aa151db0d5ded8ae3293ad9b9552e

Observation 81a29911-f563-4fdb-99d5-79ca2eeb6d7f · outbound

This paper cites Strength of minibatch noise in SGD.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Strength of minibatch noise in SGD

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.465143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.569826Z digest=sha256:6c3f2c0ef6ab411a34fd954b87c1ce3cce32fac5f7d573ead31f4ef1312fb18a

Observation 23174278-5a2e-4102-a24b-ad149a96c37a · outbound

This paper cites Stochastic gradient descent opti- mizes over-parameterized deep ReLU networks, 2018.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic gradient descent opti- mizes over-parameterized deep ReLU networks, 2018

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:51.044331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.709225Z digest=sha256:2084036e7131ba6b3f71b508e31e6f7a99e22b2fdf20b2c63b0f1ab4d3afa526

Observation 81c054d8-a7b3-4c13-be06-505ae14b05ce · outbound

This paper cites g., unit sphere in our case), can be interpreted as fixed volume.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training g., unit sphere in our case), can be interpreted as fixed volume

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.794185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.857082Z digest=sha256:0937e6cd3c91c5745b2d0ce3d88e72690e229e782f7abbcd11d4a9daf902c1ef

Observation 93744545-3594-4ef4-80f8-fe9de8266bcb · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:50.601439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:45.936357Z digest=sha256:4528c11c7f80a0bd251bfdd1b8103353039dcbc6a96d33b27e6724f70012bd92

Observation 72fb72a9-53b7-49f6-8bd8-c24002a51a42 · outbound

This paper cites An additional justification for using the Helmholtz free energy arises from the stationary distributions of SGD.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training An additional justification for using the Helmholtz free energy arises from the stationary distributions of SGD

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:50.308842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:46.143996Z digest=sha256:9e1248b2f2af92b29328699424d87e45889228edcae9161caf78a8a062de6616

Observation 2e36e8c0-bfb9-47c7-831e-035b7aa6f484 · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:38.550624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:38.550624Z digest=sha256:8ec862747a333242f84621f7e929ba8665e07a344ff1dd6e4dc2d6a6656e6d02

Observation 11e721c2-df99-4bf3-a9f3-95e2df384f1c · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 2010

Resolution
verified exact
doi, observed 2026-08-07T12:48:46.474308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:44.474011Z digest=sha256:5e48abb03ec4062eb9d33241a6bfa5f85cd1e08af3988c95c8fcce33cec7a9e0

Observation 8b7de796-ebd5-4559-89b7-2c70f9b31ada · outbound

This paper cites Deep Learning and the Information Bottleneck Principle.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep Learning and the Information Bottleneck Principle

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:44.291906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:44.291906Z digest=sha256:f75c6256ca3f1be00f6fd61668ffcefacbb8e8fa53d6b1e9c0fb25df10c06765

Observation fc848c5b-e058-4feb-9c30-8229cef5be8e · outbound

This paper cites Essentially No Barriers in Neural Network Energy Landscape.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Essentially No Barriers in Neural Network Energy Landscape

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:38.980064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:38.980064Z digest=sha256:e03ae83414ac081408777b9727c36de4d0c366a4186d0367afbad7547e0940dc

Observation 541da72f-9f5e-4f2a-b0ea-81eb4a9d7495 · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:58.968877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:38.809372Z digest=sha256:d19b66219e944866a6478594067c1d23146910bffa24a1d57ca8c22c16ca2330

Observation ed43d29b-cfc1-4ed2-a31a-9f7e8047b64f · outbound

This paper cites an unresolved cited work.

SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:54.317558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:48:43.792243Z digest=sha256:0eecfd83256d6186afa9068408af034699106eea0db50aa9a34d75467b77edca

Pith citing papers

No inbound Pith citation observations are available.