Pith. sign in

Paper Citation Record · LEDGER

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.22578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22578 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:00.259466Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf65120c-5e39-42e5-b796-7692035685d2 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, and Ruosong Wang.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Wei Hu, Zhiyuan Li, and Ruosong Wang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:09.203822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.332413Z digest=sha256:934671a7b81818bdcf1d219c59484d97bb68b532b7317e682913590f8095853c

Observation 83dd3c12-b08d-4f3a-81a2-57cb4c270077 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.939482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.377545Z digest=sha256:b3d073d6088f5904fe29056b0d75a2d376674f6580f310bf468b120d46efe885

Observation 74c7d6ca-ef8d-41ef-a13f-d5ba1e3ad86e · outbound

This paper cites Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.668772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.451054Z digest=sha256:11d1e4c6eb12c6de3b79e6e279b48e60b4804013d913050f5d19d55634f49f35

Observation 9c1ef4dc-9322-4aff-a239-1c5c4fb6f1ca · outbound

This paper cites Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T13:13:01.118361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.531975Z digest=sha256:956ad5537d081da11ada17cb5d647eaf711ebb8f6a10aadd5b73c3156eb334c8

Observation f6a9376d-c56c-4135-ae90-47898d9cfd9c · outbound

This paper cites Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.496137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.590775Z digest=sha256:069ee2ec91e629e5e9da7fb3e2ca9cd21bff066a16dbbdd0097ad00241173c8e

Observation 792d70e4-35af-48d8-be78-090179834cc0 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.357491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.664933Z digest=sha256:807eaef386635b74c6b365db57de8c1ec26cb97d995b7b6e22c059827530c8a8

Observation 8f0db553-e388-4387-8853-a948e9210d09 · outbound

This paper cites How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.217059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.756758Z digest=sha256:e7fc4a2b8adb7fa39aedf320e25313cfb4cad3d75eb9f3b95c3f8c669d939138

Observation 9b1854d3-34d3-4136-9f96-688a9f2403c1 · outbound

This paper cites Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.843877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.914089Z digest=sha256:6a3d25a3f1f828f7b0b4ecce159e8fc0c9bb8225c13560c094030ed702d2dbdb

Observation 03f525e9-9160-44a3-a045-08d516f2efe6 · outbound

This paper cites Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.001695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.032252Z digest=sha256:510d6ea61e04ca18e1f0510d361f4f398d15c14c548fdfb86c16421847aeddb0

Observation 071cdf86-ea16-4c38-a119-33a3a0a780ff · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.949186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.152527Z digest=sha256:c0bef93e3da7c058cdedd91e7389c486599f824ad82d164b9b853aa9e6a59240

Observation 38f3a440-21cb-4a86-8869-33e040ebb0ae · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.918542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.306196Z digest=sha256:e262042fb7a7b36325403be6db5b425cc82a85771fd043768b0003d82b418698

Observation 64cdbab3-38e7-43c0-a369-5bd2ede0cb6c · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.701863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.474552Z digest=sha256:57a761967bc37b3ab085cbb164739bf86b7017e302db4aa4e694f570f762c5f6

Observation e30d8cd1-5d7c-4bdb-91ac-65168652e6dd · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.590174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.590174Z digest=sha256:a50912aa5e1c871758832adc109365425c2991732c29062ac1cbbe4a5ce62484

Observation b4bd2716-0494-4d02-bfb8-c1da6de70724 · outbound

This paper cites Bach, and Loucas Pillaud - Vivien.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bach, and Loucas Pillaud - Vivien

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.630343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.704138Z digest=sha256:ff43f7c2880dc0fb3bf674e7a20da08746abc6b2738a5fbfd2987575c8e6df84

Observation e8e52731-abf1-486f-b05d-36faeb95cf96 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.503025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.846688Z digest=sha256:cc7426b5db86ed5be5132dbc08bd3b30f72be88af281f82d8e8f68a78f30c0a5

Observation 79819969-13bd-4fc5-b1f9-a5e6b3778476 · outbound

This paper cites Kakade, and Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Kakade, and Jason D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.969941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.969941Z digest=sha256:a89297c5a659a66bcd62dde3e543372e519ab670d60cce0440ccbf56b650ce44

Observation 174a3a10-6587-46e7-8fa9-61538a6fc583 · outbound

This paper cites https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:07.256033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.058268Z digest=sha256:2f8be3eb56b35f09465ff7bf278daa2650c74876981aa714536c63fd4f2dcacc

Observation 28dc0c1e-59b1-41da-9bd7-e02db3e4f05f · outbound

This paper cites Hamprecht.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Hamprecht

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.951243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.230864Z digest=sha256:56292552bbe5a252d9f4dbcfe1c6d55a2a822040571d2029c12a6aa28b3b5279

Observation 2e16810a-a1fd-4a45-a54e-25e522cbd06a · outbound

This paper cites Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.708498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.343084Z digest=sha256:031c56cbcac467e0631c2ade5d7ceca80fdc255782d23c722e24647e84438efe

Observation 0b01b0f9-19ea-4923-afb6-ec89ca45688d · outbound

This paper cites Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.540935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.443251Z digest=sha256:2ccb0488db419a727177aade9d97d976e720afa634d8dae0f977bec5b9542a37

Observation b7873b5f-3ba9-4ee0-8a70-c6cd7daa7e7c · outbound

This paper cites An introduction to probability theory and its applications, volume 2.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization An introduction to probability theory and its applications, volume 2

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.364081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.535095Z digest=sha256:37244e3f91b2ac8253e9bcb92bcb4e22e5def91519003d0619af34b38e1e61bc

Observation 51ada76a-9c7e-49b9-b168-d8626d6c5b38 · outbound

This paper cites Vetrov, and Andrew Gordon Wilson.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Vetrov, and Andrew Gordon Wilson

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.086766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.639192Z digest=sha256:af964c25781efac014cea46e8e695c616ed38a172788b714eb9167d1e21127e8

Observation f8d588e8-3ef4-433d-ae93-2cedfb5b5b9a · outbound

This paper cites https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.911555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.757397Z digest=sha256:07608209cbeef9e1a19512b22214e31c988b012f6312757381dec6156c47f48b

Observation 28a6c832-38c5-4355-9f20-5bb655161c75 · outbound

This paper cites Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.787591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.826153Z digest=sha256:1149c35960ece9d518490b4f451480131595fa9b79dcaefadcabf87232beadde

Observation febe0f18-62a5-4f05-a621-8c1522f71328 · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.938493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:56.938493Z digest=sha256:fdc97686efee02362157ec20bb4d61f00de1a9350f5510259e9e363150011589

Observation a96253ce-1c46-4c1f-b9cc-8169e7197c6b · outbound

This paper cites Haeffele and Ren \' e Vidal.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Haeffele and Ren \' e Vidal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.023287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.023287Z digest=sha256:84acd73420f59b9988fad58067b54163e25566851ecabca754ab2321a562baf4

Observation 12dd8075-579f-4650-b68b-bc03a8f58350 · outbound

This paper cites Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.676439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.197644Z digest=sha256:23c9f3a573c0a24dedd08e931c396a835be65c358118b662a6d307ed1048cd70

Observation c6928d55-1f75-46db-bb7c-805e90b5bf41 · outbound

This paper cites Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.324053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.324053Z digest=sha256:4e3c51d8f88eaff5306e088986b7f22e6c778f890cca714d5417181281814551

Observation 752329bc-34da-4c9c-97ac-02aaa5a0259e · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.456856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.442933Z digest=sha256:4375d1acd117002892e8575ae30c261e93879596ac83642576f69d151ccf84e1

Observation 16283734-4f5a-42f1-bc36-49195400c4fa · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.560687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.560687Z digest=sha256:c351b8262c8a88e036b1c766e01e919330fd9edec7c1f3b0beab6b1229d83d44

Observation f4c597be-e459-448e-9c63-6d5b0870e528 · outbound

This paper cites Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.283568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.708312Z digest=sha256:cfc13edd4543accc0862bb48307b5602e58315837024bf3bd2b18807014488c4

Observation 46685940-4451-4b0e-8239-74ac682d9155 · outbound

This paper cites Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.082243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.835214Z digest=sha256:8655557edae4206bd8125e9522e2e44fbe3a7403dbd7439a0643c3e003e02d82

Observation eee97ce2-6b4e-4dad-983b-60b3b5db89fb · outbound

This paper cites Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.832441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.920365Z digest=sha256:c136a13a000c4d28a2775b0bd0a1cdff92c5d5ee4e3277771a524aab41b6dc4f

Observation 2dc68690-56d3-48db-8bf9-d4c0fb558842 · outbound

This paper cites Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.610229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.982098Z digest=sha256:2cabdb1a7586dc4d71ea4bfbbed3d5ec6d94eba2e804e21ce9e3b5a7ffe0e1d6

Observation e05247a5-00b8-498d-a4fc-57a01db631b3 · outbound

This paper cites Michaud, and Max Tegmark.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Michaud, and Max Tegmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.456563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.075285Z digest=sha256:4268e8fee5f7c09c193510a54b86398fe3d388c2abcd4dd7806985fe567c4bc1

Observation 79238af5-cf16-4216-898e-969faaaa50b1 · outbound

This paper cites Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.326855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.144580Z digest=sha256:97631b25d275e37e1663a8eb55431fa766f1353252e39326a49c840292c37334

Observation 7dcba60d-5130-4eb8-a06f-75bed81ef595 · outbound

This paper cites Lee, and Wei Hu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Lee, and Wei Hu

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.151036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.265403Z digest=sha256:9058008a3d660cfc98dc75caefc405fe3f2b0c631affd2709eb1b67de273bb5c

Observation 7f4f4e97-1a83-4c3e-95aa-f8f1cb223d96 · outbound

This paper cites Gradient Descent Quantizes ReLU Network Features.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent Quantizes ReLU Network Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.348208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.348208Z digest=sha256:31174eb3a6b4499c8e597fef67c6ff236a3d961e30d6a31ff33506339aafdb01

Observation bb55700f-f0c9-4f44-a56d-d5e518060c56 · outbound

This paper cites A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.426905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.426905Z digest=sha256:8aa606a3dfff76590c60a3c0799b101034453c1d6ab42873398340f7db5667e1

Observation 349c5a1b-fa92-42ed-b34e-603a66389385 · outbound

This paper cites Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.976474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.544465Z digest=sha256:ec1dd5e3444be5716463863fa5952cbc3d12fdf919ed45bd52d00353a47140af

Observation 279e937b-4b17-45e4-bad8-8d78436d72b9 · outbound

This paper cites Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.780828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.664378Z digest=sha256:cf2d6d9d7a4ec121f10cd8c1f99f6844151b09190b14887cf5579032bb73b7ab

Observation 14ad19bc-d854-492b-9efd-5753fc7dc9e5 · outbound

This paper cites In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.738470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.738470Z digest=sha256:a12f57a6a28fba04d0eb286032beefd064270d368a9a5d64603241d3591c7d94

Observation 7b1bb096-e0cc-491c-b36e-c123b365e76b · outbound

This paper cites On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.632268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.852854Z digest=sha256:57d8d2571e76cda51dd8405136fe5123c042f26fbbe97cbba76812967fa263ef

Observation 9331022f-4f6d-48ce-b281-a344eee577ab · outbound

This paper cites A Note on Connectivity of Sublevel Sets in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A Note on Connectivity of Sublevel Sets in Deep Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.964229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.964229Z digest=sha256:3e813ba01605e49610fb7f66b8d55262ee2867a4a43e9b1a02213ab22e2721da

Observation 0d225fce-7953-4a55-ae08-b58b9f4cdc38 · outbound

This paper cites When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.463658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.027410Z digest=sha256:7a2d08c121494149e7db11c8635c799bb367cecacd854d2bb46684aceddbf0e3

Observation ded24de2-dbb3-4229-b5a3-351500f6c86a · outbound

This paper cites Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:13:01.452769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.146372Z digest=sha256:d4f6d94c1b5fce147d06c800715e6db97c3a06e673b3722cb83b9d63df084280

Observation 027e40ea-cedb-4cb4-bbe5-9e2f036d8106 · outbound

This paper cites Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.250869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.231078Z digest=sha256:6cf09ebbfcfe1a1744ee97f91a18dedd867112f17f9b2d5ab27cc5af1cf1f0af

Observation af995fc2-0e66-43c6-827f-2fdf0ae520e5 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.316611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.316611Z digest=sha256:8d163025df6e290651149b68e66fd6c8fd5db46ba3dc202470b481ada6d6513c

Observation 147fe6d0-4954-400d-b601-98d8cdd4522c · outbound

This paper cites Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.419770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.419770Z digest=sha256:052adccef84aff25f1c4a8e487d4d665994748c316e5c1fd4b5df2d6b6617d3d

Observation a7b54f09-d21b-4f87-9ddd-80e61a9a7a81 · outbound

This paper cites Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.111915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.498281Z digest=sha256:8fdb233c858fbad6d3836b418b58fe4aa04d32a067b347db92544da6974283c0

Observation f60e7662-4104-42ba-b431-718fd99987c1 · outbound

This paper cites How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.939431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.586553Z digest=sha256:0965734bf5f68042471bbbf289adf86c9f0fdd704cf81eb2589992e0a49c7de0

Observation 52ee293f-9430-491e-915b-45413cc09286 · outbound

This paper cites Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.653892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.653892Z digest=sha256:d88fcd0de4c8ec9cb3e610853cbdc52eb0be1d23a39f1ac3a3db0f62a7bc0f68

Observation 3a176d66-6b8b-45b8-83d3-617a6339a7b5 · outbound

This paper cites Jamaloddin Golestani.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Jamaloddin Golestani

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.705425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.707457Z digest=sha256:3c1bf3158b28dd3a56866b6a57cb0b6ae02948ef55270143da60f02b3fbb2c2d

Observation 0855d4a2-ac27-4f9b-bfdc-55c07735c983 · outbound

This paper cites Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.467783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.787288Z digest=sha256:d9c3fad6749775668d95a81ac156a424a6a64dd67b5073ad0aa183e396570977

Observation 8fc4d611-2594-4d34-ab6c-1d366d687115 · outbound

This paper cites The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.866232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.866232Z digest=sha256:1076851967057565b2b247c687eff903ce9028b23a9e27d74552d9fa15a60c9a

Observation 42386525-5e0c-41b5-b985-efdf1ddaf1e1 · outbound

This paper cites Bandeira, and Joan Bruna.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bandeira, and Joan Bruna

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.276608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.972636Z digest=sha256:7d30fbe8befb7407209b5ea2bd6f50697b0bb55ad6f8323f7c21870e94a8f5da

Observation c461e396-90eb-48de-954c-3a24e8ad053d · outbound

This paper cites The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.091759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.068517Z digest=sha256:d3e3940d8ebcfe1d68deedf586303fb61f0a867b97162a520205296e73a5d728

Observation 064d0a4b-9a0b-4ce9-a5ac-6bb0e04fee23 · outbound

This paper cites On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:00.129590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:13:00.129590Z digest=sha256:f7ae7b884fe6394570425c9607055eb20cfd5a2322ac3bc8eca8dab10ba393ff

Observation 047842c8-cae0-4b6f-88e8-63aebf6436ef · outbound

This paper cites Woodworth, Suriya Gunasekar, Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Woodworth, Suriya Gunasekar, Jason D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.882611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.191389Z digest=sha256:884a0561049057eb80c1638aa0c11df9ae4d9a87862ae520359c0ba4863559ae

Observation a1e10686-1355-4d63-8293-7c13b534aa61 · outbound

This paper cites Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.699375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.259466Z digest=sha256:f536def736b09a9a43c979f1e9b90b2b16cdf02e494e68c9cb4176bccb9eb78e

Pith citing papers

No inbound Pith citation observations are available.