Pith. sign in

Paper Citation Record · LEDGER

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2606.04476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04476 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:02:26.496063Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57185953-a4df-4a34-9bd0-a1b5c1925799 · outbound

This paper cites Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:7e30d404e138da8a7c20c895fe840c054f4d0ee4e70bcbaa9409d8024663ae4f

Observation 9e95d3e5-d7e5-4128-93dc-b149ac67772a · outbound

This paper cites High-dimensional asymptotics of feature learning: How one gradient step improves the representation.Advances in Neural Information Processing Systems, 35:37932–37946, 2022.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks High-dimensional asymptotics of feature learning: How one gradient step improves the representation.Advances in Neural Information Processing Systems, 35:37932–37946, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:47c9af4dacf58f816a0ac15c3ed519849ca29d06282fb4cb064b607fa26c02ad

Observation 5af30a62-6788-470a-9ebb-ac848617fadf · outbound

This paper cites Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58, 1989.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58, 1989

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:d5539a9fcb46ad3f1a39e2964f6916565e304845e5f61d8ca85b86afb8a76709

Observation a06d5dc0-6eb1-49a4-95b3-36611bbc7553 · outbound

This paper cites Simplicity bias and optimization threshold in two-layer relu networks, 2025.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Simplicity bias and optimization threshold in two-layer relu networks, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:490a2bc550bd6d514b1f69c4a4307bd9b88521a10730b3ff35a3053b7cec81b2

Observation fd2508e3-f5b7-4d29-b010-934b8c2201c5 · outbound

This paper cites A Bennett concentration inequality and its application to suprema of empirical processes.C.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A Bennett concentration inequality and its application to suprema of empirical processes.C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:7dd6196399ed88308a821f2975d47ba34697df8f78d47075c5fcfe101d3bed9d

Observation 9f293b8b-51e8-40e2-a300-fb66cab3cd4d · outbound

This paper cites Globally optimal gradient descent for a convnet with gaussian inputs, 2017.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Globally optimal gradient descent for a convnet with gaussian inputs, 2017

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:cb40cfa019f8e2062d0a363e831b8e2e8d809bdbfd1deae55b150c1daadfad52

Observation 4b939cc3-a74d-4ed7-9739-dd88efbc24e1 · outbound

This paper cites Candès, Xiaodong Li, and Mahdi Soltanolkotabi.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Candès, Xiaodong Li, and Mahdi Soltanolkotabi

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:4d1730b4330a9f09e30c37f51210a09b9385916323aa3424b0fe7489229c99c6

Observation 980bddcc-1df7-428d-a3d5-501c69a7db43 · outbound

This paper cites Nonconvex rectangular matrix completion via gradient descent without ℓ2,∞regularization.IEEE Trans.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Nonconvex rectangular matrix completion via gradient descent without ℓ2,∞regularization.IEEE Trans

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:fd4004178db0ed36848459a7a2097e4a3c8ac19e517511033e6e9732afb14de3

Observation 246ee337-c9c0-4f73-a3cb-d275e5d9e9ac · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:3de2f43050cc5083862f2ebd2a1996f254ccea6e65f982a1f6a2d5abbfa70f05

Observation 495f760b-4bd6-4dc8-80b6-60276cb720f2 · outbound

This paper cites Learning a neuron by a shallow relu network: Dynamics and implicit bias for correlated inputs, 2023.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning a neuron by a shallow relu network: Dynamics and implicit bias for correlated inputs, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:ac2c3a895d93724ab80f2646b89b00e5cac4787b55383bec86c5fc4be4916b07

Observation 3dcff423-d702-40bf-a2b4-6dc350c73534 · outbound

This paper cites On lazy training in differentiable programming.Advances in Neural Information Processing Systems, 32:2937–2947, 2019.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks On lazy training in differentiable programming.Advances in Neural Information Processing Systems, 32:2937–2947, 2019

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:e22b8659ba88af0687ab09d651a6072c09fd91448ceded443983bd5c556f7336

Observation 8a8ea038-8389-45e7-ad9b-95541430fdfa · outbound

This paper cites Neural networks can learn representations with gradient descent.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural networks can learn representations with gradient descent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:2f2001af64f0611bbca70f7d8979e18961b8dc0b0a1d031904b2ae711a2cb856

Observation 6455b47c-e0a3-41c0-afac-781577401f66 · outbound

This paper cites Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity.Advances in neural information processing systems, 29, 2016.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity.Advances in neural information processing systems, 29, 2016

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:b669febc728f66c7e40b0069121e65c314dfdaf7a174961048e91eb3a8046db1

Observation 3a49fd5b-e0c6-4a37-a088-4249ba4e696a · outbound

This paper cites Gradient descent finds global minima of deep neural networks.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Gradient descent finds global minima of deep neural networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:f8591a80b26dc32665aaa336e3029ce038577db17c3d57e8755451eefe225e25

Observation 6100d680-3410-4d10-9465-618c5a1d845d · outbound

This paper cites Gradient Descent Provably Optimizes Over-parameterized Neural Networks.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Gradient Descent Provably Optimizes Over-parameterized Neural Networks

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:16:44.874627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:03dd68a7755d794580e01e09c64feb2f70b4b41c0ca6fac7f83cd55fce4d4c9d

Observation 71eaa516-27f7-4e67-9354-0dc33d66e49b · outbound

This paper cites Humus-net: Hybrid unrolled multi-scale network architecture for accelerated mri reconstruction.Advances in Neural Information Processing Systems, 35:25306–25319, 2022.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Humus-net: Hybrid unrolled multi-scale network architecture for accelerated mri reconstruction.Advances in Neural Information Processing Systems, 35:25306–25319, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:1d37c0bf862c575940662566ae733b25ffb677870526233c7a1ec167c3d827d5

Observation dba5a5cb-020e-4e10-a5af-4802b8957a44 · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:fb4c3dfe0fb7fe9f7379e303f0e9c630563746e20879af60be4cd847a4f7de61

Observation 3181f5c1-d51c-4561-a42a-a51b3347f1a0 · outbound

This paper cites Matrix completion has no spurious local minimum.Advances in Neural Information Processing Systems, 29:2973–2981, 2016.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Matrix completion has no spurious local minimum.Advances in Neural Information Processing Systems, 29:2973–2981, 2016

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:9b18d253a3219bc5cce07872ef9f07f25489d058fabf51c67f708977341205a6

Observation 72ce50a2-e2d9-40fa-89d7-0a8d40216066 · outbound

This paper cites When Do Neural Networks Outperform Kernel Methods?.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks When Do Neural Networks Outperform Kernel Methods?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.878394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:bb5aab58ad32ef4fcb2aa9d40a6877ffa4ec65341af107c6dfd7c8f2eb077ba5

Observation 0dc9521a-b450-4cc3-85f0-f74b54d3421f · outbound

This paper cites Phase retrieval under a generative prior.Advances in Neural Information Processing Systems, 31, 2018.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Phase retrieval under a generative prior.Advances in Neural Information Processing Systems, 31, 2018

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:2e0efb8113bc92d1b73add2d3ecd5542fc42eb2c0be43939d37362a496f4b4f9

Observation e4d2285b-ddad-4789-87e9-e14b879a8ab2 · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural tangent kernel: Convergence and generalization in neural networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:dd10f616d6c0b74a1f9b640a521e37820bfdf281e072cacfff50c6dc6b5b1532

Observation 8f27dee3-d753-4c12-9190-cec6ed2a350d · outbound

This paper cites Gradient descent aligns the layers of deep linear networks, 2019.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Gradient descent aligns the layers of deep linear networks, 2019

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:647390713672d1359284cc5037e24fdb1bd30815f7551c47ed51a06846829110

Observation 7684b687-cc64-4410-a15b-92adeb0fb4ed · outbound

This paper cites Kakade, and Michael I.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Kakade, and Michael I

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:634db5691d26058765372cec8e4a19a63c405d1b12fcedbf6fd15c0f9213d767

Observation 50d7a600-864b-4dc7-ae3c-f6a363bbc4bc · outbound

This paper cites Deep convolutional neural network for inverse problems in imaging.IEEE transactions on image processing, 26(9):4509–4522, 2017.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Deep convolutional neural network for inverse problems in imaging.IEEE transactions on image processing, 26(9):4509–4522, 2017

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:f8b55efd0bb37c9b8b28def54fe19bad4e16675761a66743a4449c16891f5199

Observation 84407dcb-c0f7-44fd-94e6-587873124502 · outbound

This paper cites Photo-realistic single image super-resolution using a generative adversarial network.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Photo-realistic single image super-resolution using a generative adversarial network

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:58099a0dfe1dac010a266c5897c402b777b21d0decfd2a8d96bc6ffa93dd96d5

Observation aba5695e-beb9-4e0d-85fa-025b9a9eb587 · outbound

This paper cites Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.864272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:0fa8d382c81870274f34ff65b0ada60a1572d26e7f1ddb5e7cf165c20f1f0e4f

Observation d685ddc0-19cf-46b0-b72f-413e7189bb81 · outbound

This paper cites Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.867756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:28e8b2f1370a0f125c9855c13f451232a15f70dd428fdd3cf14bbb53e97af4ac

Observation 41d97781-70de-458b-83df-456bfe6e1067 · outbound

This paper cites Rapid, robust, and reliable blind deconvolution via nonconvex optimization.Appl.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Rapid, robust, and reliable blind deconvolution via nonconvex optimization.Appl

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:91c25f415a058a615187f50bcbb8b503a8ce1fc012f639a0cf7f37710bbe7548

Observation ea752827-4461-4d1d-bd18-7ed2b773024a · outbound

This paper cites Regularized gradient descent: a non-convex recipe for fast joint blind deconvolution and demixing.Inf.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Regularized gradient descent: a non-convex recipe for fast joint blind deconvolution and demixing.Inf

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:3816aef43317fe6b1630678e01e5a43a78cf7d1dceab199311674a6cfe7cb667

Observation 47afa560-2e90-4270-8211-3a8e90dc06ee · outbound

This paper cites Implicit regularization in nonconvex statistical estimation: gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution.Found.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Implicit regularization in nonconvex statistical estimation: gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution.Found

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:e58743893d7f0670a4e7f370087307d68f2abf81a5317e7da869230b22e376fd

Observation 3f459658-b3f8-4aa4-88a9-7e626bcf57a2 · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:41312911437a80e074b951d712fab0e288dcb6021bb713f13ef8237f7fb23aaa

Observation fd2fd0af-e3fe-4d7c-bdb5-3776b8bf816e · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:8817059d90ef480a1a21775c970395c34a53c63fad92f67d6b0b1d94c8b3e462

Observation b635dc95-e41e-498d-be21-720f3be637fd · outbound

This paper cites A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.881944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:434bdb4d7c0b33396e7009d31634cb2b5b25857d2e9e256f0796f666f729c386

Observation f904e8fd-631a-4cb3-b7b5-acd293d1f1e8 · outbound

This paper cites Overparameterized nonlinear learning: Gradient descent takes the shortest path? InInternational Conference on Machine Learning, pages 4951–4960.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Overparameterized nonlinear learning: Gradient descent takes the shortest path? InInternational Conference on Machine Learning, pages 4951–4960

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:2d995755ce1f839b65a6b2e617e522523d287d72b005c48e24f3997233cc1331

Observation 1c3ec22f-c3a4-4e54-9ef4-1b4026774f2f · outbound

This paper cites Towards moderate overparameterization: global convergence guarantees for training shallow neural networks.IEEE Journal on Selected Areas in Information Theory, 2020.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Towards moderate overparameterization: global convergence guarantees for training shallow neural networks.IEEE Journal on Selected Areas in Information Theory, 2020

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:6e8164b4e7d06ba9436b96a9d37bdfaff91091faff6c02bab1e794ad11553bd0

Observation 0d7aad34-f164-4b07-9ed0-af519f7c0fe7 · outbound

This paper cites Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:4d8fb429d6464780aa152a042d40dbc51fbe4625c94ea61fe698f19366c3d16a

Observation 63e86a6b-9d3d-4c52-9451-a3e42b22ff0d · outbound

This paper cites Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:30213cb0eec2e7c014b7d827bceb4b721156ae90b8a3948d9511ac1d51f41cb0

Observation fbd3d35f-7476-41dc-956e-1e0e22cb58d1 · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:8c8f4ba8344580ebc5a7cd8448293e9f0d81dfa4e11c6b121d7080365c377bff

Observation 87222743-2ab5-4b19-910e-82b68662a8ec · outbound

This paper cites Learning relus via gradient descent.Advances in neural information processing systems, 30, 2017.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning relus via gradient descent.Advances in neural information processing systems, 30, 2017

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:f5d21a32f06fd92ca6856e67081f562880e79a9a5d9817afd835b0335ba994e3

Observation 7227e63c-7854-4dbc-97cd-1181472fb6a5 · outbound

This paper cites Theoretical insights into the optimization landscape of over-parameterized shallow neural networks.IEEE Transactions on Information Theory, 65(2):742–769, 2018.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Theoretical insights into the optimization landscape of over-parameterized shallow neural networks.IEEE Transactions on Information Theory, 65(2):742–769, 2018

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:264519f37aa4bb34694dff7dfad4eb9f1f81360e66d27823779aca71e93fb5d9

Observation 756b71e2-3a6b-405b-854c-1dda97b2b8a3 · outbound

This paper cites Implicit balancing and regularization: Generalization and convergence guarantees for overparameterized asymmetric matrix sensing.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Implicit balancing and regularization: Generalization and convergence guarantees for overparameterized asymmetric matrix sensing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:48b7e068061ee0292c4c8db30dcc0baa6993845a5a9eb52a4285b19cea3eae2e

Observation 78b89da2-4051-4c69-aeb5-2ae94f44dbe0 · outbound

This paper cites End-to-end variational networks for accelerated mri reconstruction.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks End-to-end variational networks for accelerated mri reconstruction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:598d2b1fde06a611d9149938528837c2fce5766644cb359cc84fa842419c7990

Observation ebc90a49-d7cf-457a-9d41-fd359d60b2c1 · outbound

This paper cites Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:1593606c0a78fe1e84616004589face80b6a04edbb2ee96431271e62c0b07a99

Observation 62a2fc1b-a708-4ff5-9826-387102365028 · outbound

This paper cites When Are Nonconvex Problems Not Scary?.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks When Are Nonconvex Problems Not Scary?

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:16:44.871423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:feb1d57d9fa5d633dc6f3cc0d435cff012e90a11c9a6a984f088341a6a09b3bd

Observation 61f1fb50-214e-4390-8ba2-fb0daddf41ab · outbound

This paper cites A geometric analysis of phase retrieval.Found.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A geometric analysis of phase retrieval.Found

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:30b8ca29ac245e02dfc395209697f7daf27664b13322c01bd4dc16617a119a29

Observation d35d7fff-b8d6-4bb9-b567-9c4c4fe74d5f · outbound

This paper cites Low-rank solutions of linear matrix equations via procrustes flow.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Low-rank solutions of linear matrix equations via procrustes flow

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:d2e742a48bd0e2c4adb77c5bade5ca78732cfbc12037ebf3ba3b8257347db382

Observation ddc13a1c-d38b-4414-998e-18e57f19c474 · outbound

This paper cites Van der Vaart and Jon A.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Van der Vaart and Jon A

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:cb8e7c2ad4c3f4f5d4b98b7a9a3f4e6f5fc4aa1ec664fe79f5752b616df6aeb8

Observation e375c9c7-d574-4f1f-aad7-4ed5bbecab58 · outbound

This paper cites Learning a single neuron with bias using gradient descent, 2022.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning a single neuron with bias using gradient descent, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:9e9279b84659dba730661f0b473257413a5f55b32c2f62a06bd487f2023368f1

Observation dcdc7dde-9c4c-467c-864e-1792469da515 · outbound

This paper cites Wainwright.High-dimensional statistics.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Wainwright.High-dimensional statistics

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:5b37ba4e596f13edde2cf719dd6ba63fe6f9b5edb3a89b2ab9fd6d4678ef348d

Observation da3bdf4a-ac7c-46dc-b313-2006c733f1dd · outbound

This paper cites Image inpainting via generative multi-column convolutional neural networks.Advances in neural information processing systems, 31, 2018.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Image inpainting via generative multi-column convolutional neural networks.Advances in neural information processing systems, 31, 2018

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:e1777309908de8079aa37dd535bb3a9cc31c3d47d31eaeddd8382bbb1dbbc769

Observation 8c82d051-1ac8-4237-b812-fa624d2c437f · outbound

This paper cites Large learning rate tames homogeneity: Convergence and balancing effect, 2022.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Large learning rate tames homogeneity: Convergence and balancing effect, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:a719464fcca3825ae69ee97f54bc2d1816a95204b70adb9c605fb1f34a61f8a7

Observation 81fac2b4-4cf9-49ca-8370-8169aea3f154 · outbound

This paper cites Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult, 2023.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:ff7df529a01c934d0bc7fd42bd95af2174fcbd57059d984b74b40734ff77f4bf

Observation e96f5f4a-67f1-4a56-a317-602b0775b5f5 · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:f9cece3439d4aacf1863de690915a970ef0b150cda274e4570e4f4494ea71422

Observation 596557ce-5e47-4b9c-9d62-32156b5a4958 · outbound

This paper cites an unresolved cited work.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:04e378c03d39d8fd64e3f6d37c4333dc41ce1208335c0e3f70750ed8e09558e2

Observation 19f688ef-6516-4f17-975e-198422718046 · outbound

This paper cites Learning a single neuron with gradient methods, 2022.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning a single neuron with gradient methods, 2022

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:7e9f571ed7b38302c08c1efbee80560edcf8091f858165367217c552b1de40b6

Observation fda81545-33ef-4469-9b13-203c01e3a889 · outbound

This paper cites Zhang, Somayeh Sojoudi, and Javad Lavaei.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Zhang, Somayeh Sojoudi, and Javad Lavaei

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:3c1b44b9d58d8d1d932f3ae2a50c2dc6a5ce8e05847dedd0b9f6be7eb80d558b

Observation 2a55c107-15df-49cf-b214-9510fd40dc36 · outbound

This paper cites Learning one-hidden-layer relu networks via gradient descent, 2018.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Learning one-hidden-layer relu networks via gradient descent, 2018

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:8ff5a1b6bca9207e711ba52b35ab028be36f0aff5e666e14839481a927195781

Observation 359292a2-35e6-4976-b1ae-547ebe25bda5 · outbound

This paper cites A hitting time analysis of stochastic gradient langevin dynamics.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks A hitting time analysis of stochastic gradient langevin dynamics

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:3e4bcd7c44125534a048b5f1111efe5ff8d2f7daa20f3bf5dbadb12b2736d0bb

Observation 2efa4b6f-85cc-4342-adc8-27100279b3e8 · outbound

This paper cites Bartlett, and Inderjit S.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Bartlett, and Inderjit S

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:cd5b435a36fcd08c0751c821d0b11a8def1b890ff6f4714cd9ef23cda20d5cd7

Observation a4868aed-ff9f-4e2b-aa3d-003293fc4c6d · outbound

This paper cites How gradient descent balances features: A dynamical analysis for two-layer neural networks.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks How gradient descent balances features: A dynamical analysis for two-layer neural networks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:28d7e4eaf981f09228b48ed825719aa2773325f9025f06f4558258372487eb8b

Observation f8acf33b-3d50-4302-b93e-ea5f66a40076 · outbound

This paper cites Here we use the fact that µv2 1 ≤c 0c2 2 ≤ 1.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks Here we use the fact that µv2 1 ≤c 0c2 2 ≤ 1

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:d5d228f7bf78d8318db6a8bfa027f14a2a2ff484a05ec0da8d89505a09211ef6

Observation b9b5573f-e965-42c3-b61f-02dd1da4eb8a · outbound

This paper cites As a result, the sum v1 +∥w 1∥ will increase by a factor 1 + 1 8 µ∥a∥ as long as v1 ∥w1∥< 1 4 ∥a∥.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks As a result, the sum v1 +∥w 1∥ will increase by a factor 1 + 1 8 µ∥a∥ as long as v1 ∥w1∥< 1 4 ∥a∥

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:6d53458e893124dd28702126108a85bb8a17b031c0121ac0bfc1b083eb2a1a31

Observation 8a070a42-252e-4975-b378-f648de45ba4a · outbound

This paper cites v(τ) 1 v(τ) 2 #! W (τ) −W 2 F . 45 By further upper bounding the right-hand side using the fact that the ReLU activation is 1-Lipschitz, we have L θ(T) ≤ diag.

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks v(τ) 1 v(τ) 2 #! W (τ) −W 2 F . 45 By further upper bounding the right-hand side using the fact that the ReLU activation is 1-Lipschitz, we have L θ(T) ≤ diag

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T07:02:26.496063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T07:02:26.496063Z digest=sha256:bb13505dd4442250f2344abfeae465dd05986edffdb62211955ee00b14daefac

Pith citing papers

No inbound Pith citation observations are available.