Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:02:26.496063Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2606.04476.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:02:26.496063Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57185953-a4df-4a34-9bd0-a1b5c1925799 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e95d3e5-d7e5-4128-93dc-b149ac67772a · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks High-dimensional asymptotics of feature learning: How one gradient step improves the representation.Advances in Neural Information Processing Systems, 35:37932–37946, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af30a62-6788-470a-9ebb-ac848617fadf · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58, 1989
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06d5dc0-6eb1-49a4-95b3-36611bbc7553 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Simplicity bias and optimization threshold in two-layer relu networks, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd2508e3-f5b7-4d29-b010-934b8c2201c5 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A Bennett concentration inequality and its application to suprema of empirical processes.C
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f293b8b-51e8-40e2-a300-fb66cab3cd4d · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Globally optimal gradient descent for a convnet with gaussian inputs, 2017
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b939cc3-a74d-4ed7-9739-dd88efbc24e1 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Candès, Xiaodong Li, and Mahdi Soltanolkotabi
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980bddcc-1df7-428d-a3d5-501c69a7db43 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Nonconvex rectangular matrix completion via gradient descent without ℓ2,∞regularization.IEEE Trans
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246ee337-c9c0-4f73-a3cb-d275e5d9e9ac · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495f760b-4bd6-4dc8-80b6-60276cb720f2 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning a neuron by a shallow relu network: Dynamics and implicit bias for correlated inputs, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dcff423-d702-40bf-a2b4-6dc350c73534 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks On lazy training in differentiable programming.Advances in Neural Information Processing Systems, 32:2937–2947, 2019
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8ea038-8389-45e7-ad9b-95541430fdfa · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural networks can learn representations with gradient descent
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6455b47c-e0a3-41c0-afac-781577401f66 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity.Advances in neural information processing systems, 29, 2016
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a49fd5b-e0c6-4a37-a088-4249ba4e696a · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Gradient descent finds global minima of deep neural networks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6100d680-3410-4d10-9465-618c5a1d845d · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71eaa516-27f7-4e67-9354-0dc33d66e49b · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Humus-net: Hybrid unrolled multi-scale network architecture for accelerated mri reconstruction.Advances in Neural Information Processing Systems, 35:25306–25319, 2022
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dba5a5cb-020e-4e10-a5af-4802b8957a44 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3181f5c1-d51c-4561-a42a-a51b3347f1a0 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Matrix completion has no spurious local minimum.Advances in Neural Information Processing Systems, 29:2973–2981, 2016
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ce50a2-e2d9-40fa-89d7-0a8d40216066 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks When Do Neural Networks Outperform Kernel Methods?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0dc9521a-b450-4cc3-85f0-f74b54d3421f · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Phase retrieval under a generative prior.Advances in Neural Information Processing Systems, 31, 2018
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4d2285b-ddad-4789-87e9-e14b879a8ab2 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural tangent kernel: Convergence and generalization in neural networks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f27dee3-d753-4c12-9190-cec6ed2a350d · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Gradient descent aligns the layers of deep linear networks, 2019
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7684b687-cc64-4410-a15b-92adeb0fb4ed · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Kakade, and Michael I
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d7a600-864b-4dc7-ae3c-f6a363bbc4bc · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Deep convolutional neural network for inverse problems in imaging.IEEE transactions on image processing, 26(9):4509–4522, 2017
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84407dcb-c0f7-44fd-94e6-587873124502 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Photo-realistic single image super-resolution using a generative adversarial network
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba5695e-beb9-4e0d-85fa-025b9a9eb587 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d685ddc0-19cf-46b0-b72f-413e7189bb81 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 41d97781-70de-458b-83df-456bfe6e1067 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Rapid, robust, and reliable blind deconvolution via nonconvex optimization.Appl
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea752827-4461-4d1d-bd18-7ed2b773024a · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Regularized gradient descent: a non-convex recipe for fast joint blind deconvolution and demixing.Inf
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47afa560-2e90-4270-8211-3a8e90dc06ee · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Implicit regularization in nonconvex statistical estimation: gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution.Found
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f459658-b3f8-4aa4-88a9-7e626bcf57a2 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd2fd0af-e3fe-4d7c-bdb5-3776b8bf816e · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b635dc95-e41e-498d-be21-720f3be637fd · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f904e8fd-631a-4cb3-b7b5-acd293d1f1e8 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Overparameterized nonlinear learning: Gradient descent takes the shortest path? InInternational Conference on Machine Learning, pages 4951–4960
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c3ec22f-c3a4-4e54-9ef4-1b4026774f2f · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Towards moderate overparameterization: global convergence guarantees for training shallow neural networks.IEEE Journal on Selected Areas in Information Theory, 2020
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7aad34-f164-4b07-9ed0-af519f7c0fe7 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e86a6b-9d3d-4c52-9451-a3e42b22ff0d · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd3d35f-7476-41dc-956e-1e0e22cb58d1 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87222743-2ab5-4b19-910e-82b68662a8ec · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning relus via gradient descent.Advances in neural information processing systems, 30, 2017
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7227e63c-7854-4dbc-97cd-1181472fb6a5 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Theoretical insights into the optimization landscape of over-parameterized shallow neural networks.IEEE Transactions on Information Theory, 65(2):742–769, 2018
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 756b71e2-3a6b-405b-854c-1dda97b2b8a3 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Implicit balancing and regularization: Generalization and convergence guarantees for overparameterized asymmetric matrix sensing
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b89da2-4051-4c69-aeb5-2ae94f44dbe0 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks End-to-end variational networks for accelerated mri reconstruction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc90a49-d7cf-457a-9d41-fd359d60b2c1 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a2fc1b-a708-4ff5-9826-387102365028 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks When Are Nonconvex Problems Not Scary?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61f1fb50-214e-4390-8ba2-fb0daddf41ab · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A geometric analysis of phase retrieval.Found
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d35d7fff-b8d6-4bb9-b567-9c4c4fe74d5f · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Low-rank solutions of linear matrix equations via procrustes flow
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc13a1c-d38b-4414-998e-18e57f19c474 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Van der Vaart and Jon A
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e375c9c7-d574-4f1f-aad7-4ed5bbecab58 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning a single neuron with bias using gradient descent, 2022
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdc7dde-9c4c-467c-864e-1792469da515 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Wainwright.High-dimensional statistics
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3bdf4a-ac7c-46dc-b313-2006c733f1dd · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Image inpainting via generative multi-column convolutional neural networks.Advances in neural information processing systems, 31, 2018
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c82d051-1ac8-4237-b812-fa624d2c437f · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Large learning rate tames homogeneity: Convergence and balancing effect, 2022
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81fac2b4-4cf9-49ca-8370-8169aea3f154 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult, 2023
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96f5f4a-67f1-4a56-a317-602b0775b5f5 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596557ce-5e47-4b9c-9d62-32156b5a4958 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f688ef-6516-4f17-975e-198422718046 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning a single neuron with gradient methods, 2022
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fda81545-33ef-4469-9b13-203c01e3a889 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Zhang, Somayeh Sojoudi, and Javad Lavaei
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a55c107-15df-49cf-b214-9510fd40dc36 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning one-hidden-layer relu networks via gradient descent, 2018
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359292a2-35e6-4976-b1ae-547ebe25bda5 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A hitting time analysis of stochastic gradient langevin dynamics
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efa4b6f-85cc-4342-adc8-27100279b3e8 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Bartlett, and Inderjit S
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4868aed-ff9f-4e2b-aa3d-003293fc4c6d · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks How gradient descent balances features: A dynamical analysis for two-layer neural networks
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8acf33b-3d50-4302-b93e-ea5f66a40076 · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Here we use the fact that µv2 1 ≤c 0c2 2 ≤ 1
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b5573f-e965-42c3-b61f-02dd1da4eb8a · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks As a result, the sum v1 +∥w 1∥ will increase by a factor 1 + 1 8 µ∥a∥ as long as v1 ∥w1∥< 1 4 ∥a∥
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a070a42-252e-4975-b378-f648de45ba4a · outbound
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks v(τ) 1 v(τ) 2 #! W (τ) −W 2 F . 45 By further upper bounding the right-hand side using the fact that the ReLU activation is 1-Lipschitz, we have L θ(T) ≤ diag
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.