Pith. sign in

Paper Citation Record · LEDGER

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.22578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22578 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:00.259466Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf65120c-5e39-42e5-b796-7692035685d2 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, and Ruosong Wang.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Wei Hu, Zhiyuan Li, and Ruosong Wang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:09.203822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.332413Z digest=sha256:0c2003f9be18e310b879b2da01ddaadf50677e220f5db3d46c62d55cc99b4807

Observation 83dd3c12-b08d-4f3a-81a2-57cb4c270077 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.939482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.377545Z digest=sha256:bedaf12de0003ad9f61a613a15c29ef41c710c2435f3b736c6fd4389e4738591

Observation 74c7d6ca-ef8d-41ef-a13f-d5ba1e3ad86e · outbound

This paper cites Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.668772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.451054Z digest=sha256:8d36cc9ba2586bb6c75926e2348721670039b195c2923df7ff2ea2574ed412ce

Observation 9c1ef4dc-9322-4aff-a239-1c5c4fb6f1ca · outbound

This paper cites Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T13:13:01.118361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.531975Z digest=sha256:09d7bb4544b45f978bcc69bb058d94cbaddd383ab3c32e337f3ffd4d505c67d3

Observation f6a9376d-c56c-4135-ae90-47898d9cfd9c · outbound

This paper cites Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.496137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.590775Z digest=sha256:debc88ea1def28bf40b55839f3fe61e1a835b2d4a70a43577b5986d052e959f8

Observation 792d70e4-35af-48d8-be78-090179834cc0 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.357491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.664933Z digest=sha256:ed7ed189e89e83c0533fee3add87e1fddd0838afe511cbc0a05403f76a38d002

Observation 8f0db553-e388-4387-8853-a948e9210d09 · outbound

This paper cites How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.217059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.756758Z digest=sha256:10f834358093e9651c08ec63b45a41e2a7f0c6147ef928175681087558c1fc2d

Observation 9b1854d3-34d3-4136-9f96-688a9f2403c1 · outbound

This paper cites Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.843877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.914089Z digest=sha256:eb122dd34ca7827db1675bdbe4c9de1ba0b4f8aa80ef6c631ba94978753131e3

Observation 03f525e9-9160-44a3-a045-08d516f2efe6 · outbound

This paper cites Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.001695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.032252Z digest=sha256:dfac37b446cfd1622fafd518142d522b99fd69891b292a6c941bff4a67af9312

Observation 071cdf86-ea16-4c38-a119-33a3a0a780ff · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.949186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.152527Z digest=sha256:08108c9176a3c8db09deab226ba5219f72b828af03d20e58deee3e59c7ae0857

Observation 38f3a440-21cb-4a86-8869-33e040ebb0ae · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.918542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.306196Z digest=sha256:d57fb1dcaac2c51b46e2b9827adb61b4a4fda421d3913653ad98babf2646ca5a

Observation 64cdbab3-38e7-43c0-a369-5bd2ede0cb6c · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.701863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.474552Z digest=sha256:95d1ad4c403d51cec9723c028d12e9594c32a73577e8ec5cde2f7e344c2f5ec8

Observation e30d8cd1-5d7c-4bdb-91ac-65168652e6dd · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.590174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.590174Z digest=sha256:1d967eeb745fcbde4705433422340fa493cde78dfe87de5f629ddee4837a0bd1

Observation b4bd2716-0494-4d02-bfb8-c1da6de70724 · outbound

This paper cites Bach, and Loucas Pillaud - Vivien.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bach, and Loucas Pillaud - Vivien

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.630343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.704138Z digest=sha256:3ec882fdf58fbdab27920f5c14fa14658e593c428940c526ded7763b217e5c54

Observation e8e52731-abf1-486f-b05d-36faeb95cf96 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.503025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.846688Z digest=sha256:1036fb5901bef2cbacd6d510b2133174c8caf9d63930f6b2af481b177d5fa96c

Observation 79819969-13bd-4fc5-b1f9-a5e6b3778476 · outbound

This paper cites Kakade, and Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Kakade, and Jason D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.969941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.969941Z digest=sha256:9c731eb6f5740edd320734e1a6b67713c8123f48a42b37d1beacaecfd4007f52

Observation 174a3a10-6587-46e7-8fa9-61538a6fc583 · outbound

This paper cites https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:07.256033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.058268Z digest=sha256:e748481f8299432ab5c5345594ad019b4806cfd1d6b0d635954aea20397b754a

Observation 28dc0c1e-59b1-41da-9bd7-e02db3e4f05f · outbound

This paper cites Hamprecht.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Hamprecht

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.951243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.230864Z digest=sha256:0090a2aedde3bfb78b2f38fe390b3c92f98abc4b24cb96dcefdc44e5903e862e

Observation 2e16810a-a1fd-4a45-a54e-25e522cbd06a · outbound

This paper cites Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.708498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.343084Z digest=sha256:5765c1021621dc738db4fa20a3d24d7f3d3f4a7186feeaee98e19a8eb7257c97

Observation 0b01b0f9-19ea-4923-afb6-ec89ca45688d · outbound

This paper cites Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.540935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.443251Z digest=sha256:a04fa4c80e81dd1eddacb246ab0632973c3cfb5e9685a2b25bd606015e169a04

Observation b7873b5f-3ba9-4ee0-8a70-c6cd7daa7e7c · outbound

This paper cites An introduction to probability theory and its applications, volume 2.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization An introduction to probability theory and its applications, volume 2

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.364081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.535095Z digest=sha256:2feeef769f0a99138c77580a63da143f448c25d3a88316aabcddd8286553bf6e

Observation 51ada76a-9c7e-49b9-b168-d8626d6c5b38 · outbound

This paper cites Vetrov, and Andrew Gordon Wilson.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Vetrov, and Andrew Gordon Wilson

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.086766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.639192Z digest=sha256:13039756ed689f5456f9ffa3640795a83f17161dfb2bdf38cb7f44224e3f2ff4

Observation f8d588e8-3ef4-433d-ae93-2cedfb5b5b9a · outbound

This paper cites https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.911555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.757397Z digest=sha256:d5764e690da174a120e00f84b3badba9fc1495eb9642fec9d04a68e6f804de97

Observation 28a6c832-38c5-4355-9f20-5bb655161c75 · outbound

This paper cites Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.787591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.826153Z digest=sha256:39e0d77cd7f9e37892ce59e5e279e9c48b5be073f8511f1211b3f81178500ee1

Observation febe0f18-62a5-4f05-a621-8c1522f71328 · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.938493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:56.938493Z digest=sha256:48097f031b80913a79914413bf19e6fd37f96c240906aa6f3283927797a55ab4

Observation a96253ce-1c46-4c1f-b9cc-8169e7197c6b · outbound

This paper cites Haeffele and Ren \' e Vidal.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Haeffele and Ren \' e Vidal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.023287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.023287Z digest=sha256:b751cdece35fea14660ea20a99bb1f1074660bf5a73a5228a1b878456b0a18bc

Observation 12dd8075-579f-4650-b68b-bc03a8f58350 · outbound

This paper cites Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.676439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.197644Z digest=sha256:97206a49d491e333418e2ec3b42c0c8e446abb5dd44ddde5a3d11996aa3ecf98

Observation c6928d55-1f75-46db-bb7c-805e90b5bf41 · outbound

This paper cites Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.324053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.324053Z digest=sha256:0e6ec782b3b18e7d8eda6666821aaf1cd0b85ad3f9b5c6d15c71e1a72cd36c3d

Observation 752329bc-34da-4c9c-97ac-02aaa5a0259e · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.456856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.442933Z digest=sha256:ae653ed5a98547f4acc189dc20bbd295ec8f3e03e1ca55bac6aabf3f622c4a37

Observation 16283734-4f5a-42f1-bc36-49195400c4fa · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.560687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.560687Z digest=sha256:fc4e782aea2b19cd56c31de9fc632dd311e61103e9563702e467695414101891

Observation f4c597be-e459-448e-9c63-6d5b0870e528 · outbound

This paper cites Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.283568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.708312Z digest=sha256:b17c1a9303942559764e92596446385ce3fda087acd47448c9a79405ddbc1200

Observation 46685940-4451-4b0e-8239-74ac682d9155 · outbound

This paper cites Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.082243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.835214Z digest=sha256:5aac3776df340894662410df27b3e5862f93397739c6cd8d9ad269bd1ae30eb4

Observation eee97ce2-6b4e-4dad-983b-60b3b5db89fb · outbound

This paper cites Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.832441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.920365Z digest=sha256:ae567170f29cfe31d199b9f55377174afc320eec0c5d44e1d2984a9e10bc4795

Observation 2dc68690-56d3-48db-8bf9-d4c0fb558842 · outbound

This paper cites Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.610229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.982098Z digest=sha256:80b3f91229df424e6ccab6eed28e7b4737a37024865cdf490dd68d7736993f57

Observation e05247a5-00b8-498d-a4fc-57a01db631b3 · outbound

This paper cites Michaud, and Max Tegmark.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Michaud, and Max Tegmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.456563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.075285Z digest=sha256:6e47bc4b0704068db3312aab69bf68d6206799c2b0565dd8090df33ae2de173f

Observation 79238af5-cf16-4216-898e-969faaaa50b1 · outbound

This paper cites Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.326855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.144580Z digest=sha256:3bbb8104341e74a6aec7ed5fd03bc3b402058ce705e9f09c116f9e37b7883590

Observation 7dcba60d-5130-4eb8-a06f-75bed81ef595 · outbound

This paper cites Lee, and Wei Hu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Lee, and Wei Hu

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.151036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.265403Z digest=sha256:e3eae0e6afb36d78190e6bf16f6ccf38a21b4b85713e96f8b1443c87b0668cd0

Observation 7f4f4e97-1a83-4c3e-95aa-f8f1cb223d96 · outbound

This paper cites Gradient Descent Quantizes ReLU Network Features.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent Quantizes ReLU Network Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.348208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.348208Z digest=sha256:965a763f0c87dd45012b8f026282ab82cdf143389ef3c89a46bfae46e0f56794

Observation bb55700f-f0c9-4f44-a56d-d5e518060c56 · outbound

This paper cites A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.426905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.426905Z digest=sha256:f30e293952126d251b7a0fc6109bffa6afca64402ccc5e05a8d2488511cdec78

Observation 349c5a1b-fa92-42ed-b34e-603a66389385 · outbound

This paper cites Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.976474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.544465Z digest=sha256:2dcb63199f0f208f10169a2f438984b6d9c5d39a61c047dfebbdb1aa7e82f7a4

Observation 279e937b-4b17-45e4-bad8-8d78436d72b9 · outbound

This paper cites Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.780828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.664378Z digest=sha256:0fac99aba1fe79b44ee51200e02dba8499642028241d5c0663b1af9506f95e01

Observation 14ad19bc-d854-492b-9efd-5753fc7dc9e5 · outbound

This paper cites In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.738470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.738470Z digest=sha256:06a413cdfcb694638e94288b3304f3a503210b23b0bec72c539a97229a72b67f

Observation 7b1bb096-e0cc-491c-b36e-c123b365e76b · outbound

This paper cites On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.632268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.852854Z digest=sha256:2c7e2dc433728c699d498202764bbb106d03b198f71cd213c5da94a1b8797b48

Observation 9331022f-4f6d-48ce-b281-a344eee577ab · outbound

This paper cites A Note on Connectivity of Sublevel Sets in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A Note on Connectivity of Sublevel Sets in Deep Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.964229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.964229Z digest=sha256:3e2ecbfd978207db4c879e329a881a63f08d068c018cd1c404657df7faf52c7a

Observation 0d225fce-7953-4a55-ae08-b58b9f4cdc38 · outbound

This paper cites When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.463658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.027410Z digest=sha256:947d94d463c6e1d57144e0b13c1e01268f1251602362c0f11d2deda60fee5be1

Observation ded24de2-dbb3-4229-b5a3-351500f6c86a · outbound

This paper cites Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:13:01.452769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.146372Z digest=sha256:e87096b5e547c571e24899886c1a72cfc44eab6415f53dc1d20b9490ba8c8e7b

Observation 027e40ea-cedb-4cb4-bbe5-9e2f036d8106 · outbound

This paper cites Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.250869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.231078Z digest=sha256:aa2743e794c00b790336a09b91e08ed974fbb6c5da6763a573ff8cc6a08d0f13

Observation af995fc2-0e66-43c6-827f-2fdf0ae520e5 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.316611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.316611Z digest=sha256:fc46a1704ce124f21bc811efd1c2feb1374159f84b53935c48cb6e789e4487a5

Observation 147fe6d0-4954-400d-b601-98d8cdd4522c · outbound

This paper cites Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.419770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.419770Z digest=sha256:da13e22ed5c5d8f18c3d762e6b45ff73626fc1e67a6f2f11570d37e55d867db6

Observation a7b54f09-d21b-4f87-9ddd-80e61a9a7a81 · outbound

This paper cites Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.111915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.498281Z digest=sha256:c2e09f1b8ec6db32655777225a9d8640dc90143fe86aec6f3de1ef49230db953

Observation f60e7662-4104-42ba-b431-718fd99987c1 · outbound

This paper cites How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.939431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.586553Z digest=sha256:98b19c974888341764181fec4b909f6669fb0e06ca29a25fe466595699181a30

Observation 52ee293f-9430-491e-915b-45413cc09286 · outbound

This paper cites Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.653892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.653892Z digest=sha256:fe85d9c4118a50f0548d99a5d77f52d7459a466a7ff96185da869377abefc2cb

Observation 3a176d66-6b8b-45b8-83d3-617a6339a7b5 · outbound

This paper cites Jamaloddin Golestani.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Jamaloddin Golestani

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.705425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.707457Z digest=sha256:5833d8baf7d23d9952f6c163b072ec8585312ffbe5f0edd46fe8bee3e8e57219

Observation 0855d4a2-ac27-4f9b-bfdc-55c07735c983 · outbound

This paper cites Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.467783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.787288Z digest=sha256:cff7d4acb0f183c5e0882e17d7219be4d3ab3e5f153ead86de8715494d975833

Observation 8fc4d611-2594-4d34-ab6c-1d366d687115 · outbound

This paper cites The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.866232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.866232Z digest=sha256:9fdc3d567a4ff757ff94670ce2e11a42e390dae08bcaf8af7d76160033a22817

Observation 42386525-5e0c-41b5-b985-efdf1ddaf1e1 · outbound

This paper cites Bandeira, and Joan Bruna.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bandeira, and Joan Bruna

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.276608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.972636Z digest=sha256:39a79b00774bc182a5bf0b9ffc40438b7a67efaece5ef0885aa20b027127c50b

Observation c461e396-90eb-48de-954c-3a24e8ad053d · outbound

This paper cites The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.091759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.068517Z digest=sha256:9686f8364240dad4ef42068c3774b96d7038aba3fe740f6e398b526d6d066cd9

Observation 064d0a4b-9a0b-4ce9-a5ac-6bb0e04fee23 · outbound

This paper cites On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:00.129590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:13:00.129590Z digest=sha256:da111cf330a102b2498193dca2ddc6ad0851c98a340b37c388462b9f918fb5aa

Observation 047842c8-cae0-4b6f-88e8-63aebf6436ef · outbound

This paper cites Woodworth, Suriya Gunasekar, Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Woodworth, Suriya Gunasekar, Jason D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.882611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.191389Z digest=sha256:c58280d717d028131272b7914fc660f35e58c81a06bb780b5ff168255f68a91b

Observation a1e10686-1355-4d63-8293-7c13b534aa61 · outbound

This paper cites Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.699375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.259466Z digest=sha256:fa918c6820e33fd29cd9958642617e21c423f49a3cfdca31a8aea23f81928de6

Pith citing papers

No inbound Pith citation observations are available.