Pith. sign in

Paper Citation Record · LEDGER

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

As of 10 August 2026, this Paper Citation Record lists 100 of 124 outbound references and 2 inbound Pith citation observations for arXiv:2502.01763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01763 v1

Coverage vector

measured 100 of 124 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:40.728250Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:23:02.017903Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 124 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bce02749-8550-4b65-9051-fb87ac1530e8 · outbound

This paper cites and Szepesv \'a ri, C.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Szepesv \'a ri, C

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.304467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.304467Z digest=sha256:cebec6a5a279026634e96880aec3acfba512ec6594547d1336423d1b46628bdf

Observation 7193376e-7dd6-48c3-b2a8-b33e0f8e89e2 · outbound

This paper cites B., and Misiakiewicz, T.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning B., and Misiakiewicz, T

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.309277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.309277Z digest=sha256:63ab70bb7736e7c738262b5430a85fb96dbbd982313fef82da8c41ae9e63bae2

Observation 2f320a98-16b5-4740-b412-3c65974b592d · outbound

This paper cites B., and Misiakiewicz, T.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning B., and Misiakiewicz, T

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.313480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.313480Z digest=sha256:f6c802cacd22b196cc9716e267109e69b6ed3657f8ceec5a574ec4f73cd6a081

Observation 1f82431a-f5ef-45a0-919a-6a9adada6074 · outbound

This paper cites Optimization algorithms on matrix manifolds.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Optimization algorithms on matrix manifolds

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.317319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.317319Z digest=sha256:be81492caaaf6b2a91c06c432722bfdaf66eb4ada300737932f38aeda5efc4dd

Observation de925ca4-f8c1-4d47-bbbd-fa25391ce57f · outbound

This paper cites and Pennington, J.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Pennington, J

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.321442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.321442Z digest=sha256:e95985e4daf2302e28f7c7f025fc18f1edaf5d77f16eadc0ccde37cf464b27e1

Observation fbb44cb4-3cd9-4723-9b84-6abc71b8a4a4 · outbound

This paper cites and Parrilo, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Parrilo, P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.325138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.325138Z digest=sha256:b29cbb793acb3983414ca26f64ce30b9d3164f7556a009b12227f9bb7a8be6b2

Observation bc46bb90-4b9e-4c1f-8db1-4c69c55aa12b · outbound

This paper cites When Does Preconditioning Help or Hurt Generalization?.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning When Does Preconditioning Help or Hurt Generalization?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.329559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.329559Z digest=sha256:cf282d098b2241c5c2394c19e4c1d78f5c03b011fab2eb2a5fdd5d63607998a2

Observation 4d693f72-8cd8-44b1-b08b-cd2ea91164e3 · outbound

This paper cites Locoprop: Enhancing backprop via local loss optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Locoprop: Enhancing backprop via local loss optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.334202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.334202Z digest=sha256:950a4d0e27ed9c8d1d8e7426bbc0fa7d540d8526a57833fec8daf9c2b36d8895

Observation f6ee63cb-87d5-46c8-b7e5-4bfbe809579d · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Scalable Second Order Optimization for Deep Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.338674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.338674Z digest=sha256:7ee8f95f62a7661a4ec9754c9624a3b9a0a299857910a4a9c0864ff133b40765

Observation 10a1e903-c85c-402b-b5c9-ed75f1db0ffd · outbound

This paper cites Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.343242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.343242Z digest=sha256:e7bccfa78e3c41034c2576679a6f0f1bb5c2a7750ffebc3642441e4327ef79a4

Observation ffc32a11-37db-4fb2-84a0-6b7db5fd17ad · outbound

This paper cites Implicit regularization in deep matrix factorization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Implicit regularization in deep matrix factorization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.348170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.348170Z digest=sha256:613ca514165f28935fa820fc6debf933cdfd137b145829e9961e5c8bcebfb942

Observation 9e0fea14-736a-46a7-85e0-4113327f6a78 · outbound

This paper cites B., and Martens, J.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning B., and Martens, J

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.354122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.354122Z digest=sha256:1c25efbf861d4f4b802068ff694da16d4bc92c89302f10440dd6ad365447c53a

Observation b3f1fec8-b5c5-4e6a-bedc-373c61660763 · outbound

This paper cites A., Suzuki, T., Wang, Z., Wu, D., and Yang, G.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A., Suzuki, T., Wang, Z., Wu, D., and Yang, G

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.361665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.361665Z digest=sha256:03ccf7fc4b82080cfaa0a2b47eeecf450bf2821a19815427870123d9bfc38d04

Observation b5fb6e65-fc38-4903-9e15-552dc4f0c857 · outbound

This paper cites A., Suzuki, T., Wang, Z., and Wu, D.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A., Suzuki, T., Wang, Z., and Wu, D

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.367830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.367830Z digest=sha256:325978ef9b9c0f7491ddeec2336d542912cc16e74ab5ef1d6ef55465443ba783

Observation 07808b72-662a-4e0d-aa38-38a34939e8f9 · outbound

This paper cites Scaling laws of optimization, 2024.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Scaling laws of optimization, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.372639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.372639Z digest=sha256:2a74bacadfcbc03069523fdc6ae77fd2bc0580c74b15471efd4e7fb9b53a389d

Observation 872b33da-ce0b-410a-b93f-b550286e9611 · outbound

This paper cites and Lee, J.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Lee, J

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.377075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.377075Z digest=sha256:ef13c94346366e295d4b5cc17987b541dead32acd8930afdd5543e12325b037e

Observation 67f30173-e43f-4a91-89fa-d175b3d5ef5c · outbound

This paper cites Hidden progress in deep learning: SGD learns parities near the computational limit.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Hidden progress in deep learning: SGD learns parities near the computational limit

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.381857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.381857Z digest=sha256:d37865965721e3ccf9e8edd314a486a6717e2419878722796b826a9c96d27355

Observation 298c1374-61c3-40f4-a042-331e08af5ff4 · outbound

This paper cites Online stochastic gradient descent on non-convex losses from high-dimensional inference.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Online stochastic gradient descent on non-convex losses from high-dimensional inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.386686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.386686Z digest=sha256:a11b017daf716109bbbd3d4cdac86342d996587cd1ae17d160fb9072718c6309

Observation cfa5e545-948c-440d-b037-85207d68765e · outbound

This paper cites Gradient descent on neurons and its link to approximate second-order optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Gradient descent on neurons and its link to approximate second-order optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.391577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.391577Z digest=sha256:21dde6458142cd8ed155488b57368aab27f035c0a12bfe16f011e1bb642c7e6f

Observation e104b005-4fdb-4942-9cd6-7ff6a9ed5fc9 · outbound

This paper cites Modular Duality in Deep Learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Modular Duality in Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.399031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.399031Z digest=sha256:a06244c8cadde63a11969eb7ca4caf64ba11ae0f2c323ff7bb0de65b40fbc46f

Observation b146c1de-38bf-40c4-a319-a6fd06b40fd3 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Old Optimizer, New Norm: An Anthology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.403483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.403483Z digest=sha256:f8b90e43081d143a2b7333c6a99080fa8ad1c14018f2dd569c50429defb0eb5c

Observation a64b08e4-4b17-473d-8c9e-329c965a33c7 · outbound

This paper cites Learning time-scales in two-layers neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Learning time-scales in two-layers neural networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.408891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.408891Z digest=sha256:e87038bbf447bdfb0513a511486f843b445a90df43dba568a7020815f34d9e33

Observation 1fbfaf72-6a26-49b0-88a2-c72effb17601 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.412814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.412814Z digest=sha256:2cd5d30c757456fea30ac2e48e8d6fe5bf4f81a00668fde07a96372e9ad5d1c9

Observation 30f40d95-194f-4b0e-a58f-bead24c7b0b2 · outbound

This paper cites and Mondelli, M.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Mondelli, M

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.418425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.418425Z digest=sha256:78934b41f82c69433f6ddca83b1b54f73e7226465d07d363b55b9e19ae73b233

Observation bf328dcd-c6a2-4e14-8f98-9b2938b8b042 · outbound

This paper cites Privacy for Free in the Overparameterized Regime.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Privacy for Free in the Overparameterized Regime

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-09T14:47:42.500114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.422681Z digest=sha256:2585422948c360d32ad588dddb3fdfbfb0bd62eee4ff39425222dea817a0126e

Observation 517db258-d052-4692-875b-96881237575e · outbound

This paper cites Beyond the universal law of robustness: Sharper laws for random features and neural tangent kernels.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Beyond the universal law of robustness: Sharper laws for random features and neural tangent kernels

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.428124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.428124Z digest=sha256:35dc950b2e4bd96d01c7f0371a2b1512826b54627a76d029e1fee2f4d89d8b39

Observation 8a41c11c-128f-4786-90a2-a04690a6f33b · outbound

This paper cites Practical Gauss- Netwon optimisation for deep learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Practical Gauss- Netwon optimisation for deep learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.432156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.432156Z digest=sha256:a5eef9cc20b8394c8acd0b1153fba882dc832e0e2dd20f8c3163375dad87c56d

Observation d8ade78a-5882-4f91-9a98-047061157a6e · outbound

This paper cites H., Hansen, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning H., Hansen, S

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.435829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.435829Z digest=sha256:784cf8323271b76834e027ba76282e942ec373d99237756e199692743a409ae2

Observation e9f04798-4a48-4bab-a702-aa64104d58f7 · outbound

This paper cites Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.439889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.439889Z digest=sha256:ed0c3a8464bc90b521949b54e8bcc03dccca5bb256fa6d6d0da80f0084598e73

Observation dcdce652-6755-4296-acbc-cf9716c7fb7f · outbound

This paper cites Exploiting shared representations for personalized federated learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Exploiting shared representations for personalized federated learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.443948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.443948Z digest=sha256:52399d9a75ccaa1eb3d7e80095248b377cb62cc9b3714be6a5c37ee97f825253

Observation 5d5657af-e725-4a2b-a3a8-09f1b66ba8d6 · outbound

This paper cites Provable multi-task representation learning by two-layer ReLU neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Provable multi-task representation learning by two-layer ReLU neural networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.449327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.449327Z digest=sha256:d5d60af7637e1017670fd33bcc3a6159e025e5b16da4cea3ef9691c879ddb77c

Observation 256b669c-b153-4ee5-b57c-08d152b1bdbc · outbound

This paper cites Asymptotics of feature learning in two-layer networks after one gradient-step.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Asymptotics of feature learning in two-layer networks after one gradient-step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.453841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.453841Z digest=sha256:137fc7567c5db2e350b561b1a3faa2285427f8604e1adb432d123f9a45cfd274

Observation fa88d2c3-f48e-4033-8299-440b908b08ec · outbound

This paper cites Benchmarking Neural Network Training Algorithms.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Benchmarking Neural Network Training Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.457590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.457590Z digest=sha256:bae5f9aca50bbaf070b0cae356fa38f73cec0c3b67ee372b06d31537bb5523fd

Observation b1bbd743-58fa-4b98-93d6-9628fe76a564 · outbound

This paper cites Neural networks can learn representations with gradient descent.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Neural networks can learn representations with gradient descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.462358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.462358Z digest=sha256:892e889c685a0f5b1f3729101ff35db622222a62bdc98821c6a4ddaf0f7adcdf

Observation 22cf088e-f81b-41dd-8090-046e598c859a · outbound

This paper cites How two-layer neural networks learn, one (giant) step at a time.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning How two-layer neural networks learn, one (giant) step at a time

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.466603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.466603Z digest=sha256:3ac5a9c8eb9041171813eaeb045e23b3659eac80633e61a43a7fcb8f4c471d9d

Observation 84927feb-beac-4eb3-a64a-d1fe23e2c2bf · outbound

This paper cites A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.471320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.471320Z digest=sha256:f36e535440835f47c7a3d4e1f7bfbee4cd7648da894903c64e0ae6e14d1c8979

Observation 7e01e8de-9c7e-4716-afc4-6d14356ae73b · outbound

This paper cites The benefits of reusing batches for gradient descent in two-layer networks: Breaking the curse of information and leap exponents.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The benefits of reusing batches for gradient descent in two-layer networks: Breaking the curse of information and leap exponents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.475810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.475810Z digest=sha256:e52eaff8d441c3d385ba540e82940621795b879986ab35d83d60254ca5dd0b31

Observation 73454c87-a7f4-4455-9fd0-ed35153d6aec · outbound

This paper cites Modular block-diagonal curvature approximations for feedforward architectures.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Modular block-diagonal curvature approximations for feedforward architectures

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.480451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.480451Z digest=sha256:11c3dc78ebd9331772bc3502522b1bf136517fcdb455b1b748bbcd40402932a6

Observation 6866259e-82f7-440b-84eb-964097cfc13f · outbound

This paper cites Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.486358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.486358Z digest=sha256:9657149d42976674e7402563d19f7966fa666ebfdc102c4bf2b73c95513908b5

Observation c762e9fa-410a-40ba-8934-edbd513acfc9 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.490862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.490862Z digest=sha256:86b485ed79b6b900b02dfb0156c630401272fe6cdf99c16e013274c18c3bacb1

Observation 25d7dbda-a607-4810-85b7-8cb65679855e · outbound

This paper cites and Wager, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Wager, S

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.495605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.495605Z digest=sha256:5d134147660fcdfe4d36769800a178dd3ff52d312180010fbc7bd092de19b156

Observation 15c500be-be2a-4cc5-a5e2-fa8686a15c2a · outbound

This paper cites S., Hu, W., Kakade, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning S., Hu, W., Kakade, S

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.500141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.500141Z digest=sha256:424ee24c2813273e3126cf6b5ced0ce5f87266fbdc297944e2440d3a3c13547a

Observation bc16ca31-d1cc-4744-85a8-393eb048238e · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Adaptive subgradient methods for online learning and stochastic optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.505091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.505091Z digest=sha256:d8e210b2fa448818fca3b8d286475075b743c873287710fb87bef5044f678129

Observation ff4ce0d6-8580-4bcf-8fab-76c3c9ef8204 · outbound

This paper cites Proximal backpropagation.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Proximal backpropagation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.509479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.509479Z digest=sha256:cd449011d321691031c4d67a1f5bb617b808ad6409c510fc2e9f7f05504679ec

Observation c99f2b0f-2772-4e15-9be3-6dcaa37c2bb9 · outbound

This paper cites Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T14:47:42.442595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.513091Z digest=sha256:972c39d3270110b5e34eebc288f8cdfeae816ad1f9695cd6096b2ed617089adc

Observation 252bdf19-c590-4bbc-9cd7-29e0d50eb381 · outbound

This paper cites Linearized two-layers neural networks in high dimension.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Linearized two-layers neural networks in high dimension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.516867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.516867Z digest=sha256:c11a0023130bba544be14307f2010cb9693cf72cc82cf1daf8abe95a0c010314

Observation f7931c78-ba89-4082-85ee-c5923442b1b1 · outbound

This paper cites When do neural networks outperform kernel methods? Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12), 2021 b.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning When do neural networks outperform kernel methods? Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12), 2021 b

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.520374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.520374Z digest=sha256:e10035cf4d5259e27a0c0a8015e11c275103cb4a6fadd47418de9520b02d20a3

Observation 85e99ff8-1046-4e3e-a3f8-18dd3eafd38e · outbound

This paper cites A family of variable-metric methods derived by variational means.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A family of variable-metric methods derived by variational means

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.524058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.524058Z digest=sha256:928ceb0f5f873adea0325abf6ce7f76c6cc64e7ae3ca320a8a43684839afaf4f

Observation 82c83bb6-3165-4b28-9f16-a163af501fd2 · outbound

This paper cites Practical quasi- Netwon methods for training deep neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Practical quasi- Netwon methods for training deep neural networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.528448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.528448Z digest=sha256:528282a6443a73ae4a6819d9fd69fd6133aa43bebdb931777eae402279085e3d

Observation abe35d8f-a5d0-4268-8cc0-5d98c753aa40 · outbound

This paper cites The G aussian equivalence of generative models for learning with shallow neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The G aussian equivalence of generative models for learning with shallow neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.533185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.533185Z digest=sha256:029c5e0aa76376fef964b192704a38748e5f41fc28cf6a8d32d58b08692e2fa5

Observation e457abdd-3640-4945-8bb9-6fc4db3ca515 · outbound

This paper cites Spectral Phase Transitions in Non-Linear Wigner Spiked Models.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Spectral Phase Transitions in Non-Linear Wigner Spiked Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.537284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.537284Z digest=sha256:068b9d5b62d979dd888c9699b332cf584a112142c9d1400cbeeb7c5a55a2f1b0

Observation 421b455b-ee73-471f-a4a2-ca042e11dd29 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.541398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.541398Z digest=sha256:7100defdd4730aeca1854bde0255421a6a5c9af3a18a67fdbb157a2be126e929

Observation 8126d448-d040-4a49-b6ac-1eff22992fcd · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Shampoo: Preconditioned stochastic tensor optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.545269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.545269Z digest=sha256:e343c2d33916f53314b3df71308530072982460c085338bf8df849ffef58c082

Observation bcebb231-9ae7-4250-bd93-c2e08d1f1d6d · outbound

This paper cites and Nica, M.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Nica, M

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.548962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.548962Z digest=sha256:65f8d657811efab54d1afdd4359c9b4127467734f3be8856d12d5a448bb0bc18

Observation 5a6fb108-5b8d-4047-90f3-6f33e8decf73 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.552782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.552782Z digest=sha256:7c3507d98f3b0da1044b2c7365e59203d2762b66873ef0c6250c15a84a486f75

Observation f2471c44-e407-4422-923a-84655537715f · outbound

This paper cites and Javanmard, A.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Javanmard, A

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.556647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.556647Z digest=sha256:70e0f45bddc230ca359582d05ed0dc441ad6694e7c7e062cdaa47a8ca8202816

Observation 29217597-fb3d-4d51-ae01-90ea07e95239 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.560608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.560608Z digest=sha256:b048b138d071d53c697e8cb70738053210643c04def66eee5d768426f7064aed

Observation 9c12a425-d4ce-4403-aaf6-f07e0d3354ab · outbound

This paper cites M., and Zhang, T.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning M., and Zhang, T

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.215635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.564868Z digest=sha256:6df82da54abae15cb9e76e2c0f1597d168b31580ecf7ba76de3863e919d8cb9e

Observation bc7af867-b87f-45f9-adec-00bdf6d797ae · outbound

This paper cites and Lu, Y.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Lu, Y

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.201443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.568825Z digest=sha256:f7d88a86c57cf6f90ebbdeb2d89c0a7ea276d600b9dbe81c67b6e497db4f115a

Observation f54e91de-1899-4096-adac-1fae999788d5 · outbound

This paper cites and Szegedy, C.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Szegedy, C

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.188677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.572424Z digest=sha256:6e6f2c9db17e7806af4d72b457d0ec189fd6bb260730d6e1ff22e81a149200f6

Observation f9d9e5e1-2cef-4e89-836f-81fc90171b64 · outbound

This paper cites On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.576362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.576362Z digest=sha256:34adee5ce8232a30de61a7b179e601f732330f0b154ae094a4beb801534aea59

Observation 690ad69d-fb90-4bbb-b54e-8173f53389ed · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Neural tangent kernel: Convergence and generalization in neural networks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.176504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.580605Z digest=sha256:00c247963caff31887579fdd5645ac72a9f2cb5d306a05df5f3dff57c174d85b

Observation 1279429c-1794-4140-be6d-178de71424b7 · outbound

This paper cites Low-rank matrix completion using alternating minimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Low-rank matrix completion using alternating minimization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.163081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.584397Z digest=sha256:5925f63a3ef2a258a16243be531e178cc1d7e7f5a253d36415480f2ab6f4d1af

Observation 026a4f86-c81a-4f80-9bd7-24b01f956401 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Muon: An optimizer for hidden layers in neural networks, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.147831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.589217Z digest=sha256:939db776ed8883918bbaa79846c07d6db0a9d8e72b7656243fdf27f4b8de2c26

Observation e77cab59-85e7-4356-b651-1126434c3d34 · outbound

This paper cites Improving Generalization Performance by Switching from Adam to SGD.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Improving Generalization Performance by Switching from Adam to SGD

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.593164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.593164Z digest=sha256:53a56335db0cd1daaaa7e5984c5f730cd9df5282fd5b081f811aeef8d33650b4

Observation 652e5d86-f409-41c7-ae50-3cc85dabc55a · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.597248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.597248Z digest=sha256:84f3cf2d48d88b690d32458a4edbec3118a1ddca8eae826eb532e72672899bce

Observation 869a9e7f-5c29-4588-8006-ceb7733b6b02 · outbound

This paper cites M., Ma, T., and Liang, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning M., Ma, T., and Liang, P

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.123219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.601177Z digest=sha256:8e1bcbd14677ba576dbd793148b84a73ec923b7e9e034d00743cc2af2109c31c

Observation 72e137fb-aec8-4d12-af33-abbead0af4b2 · outbound

This paper cites Scalable Optimization in the Modular Norm.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Scalable Optimization in the Modular Norm

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.604767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.604767Z digest=sha256:dc8689a15060ad190d849c2fd29909384c4a071516d260dd72ae28c9402b43c0

Observation 836c0b76-eeb2-4ac8-8076-2ae655001634 · outbound

This paper cites and Massart, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Massart, P

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.109668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.609198Z digest=sha256:08512e1a802feff14b56b857d5c9858df387a2a0b36628126db0b16531409e43

Observation f975a99a-dee8-4548-9e66-394fcce7a653 · outbound

This paper cites Demystifying disagreement-on-the-line in high dimensions.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Demystifying disagreement-on-the-line in high dimensions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.096942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.613267Z digest=sha256:3512e26aa85063898bbcd9149bfb8b3a0e0814c142f7f5cb5a10affa6a88d185

Observation 4a46ccb1-c848-4c3b-ae50-394da2df9736 · outbound

This paper cites Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.617364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.617364Z digest=sha256:d7cb6cf51669833fb4ab8c0752e3a078e9ee943cfecd65b169dfa71f1fd229bd

Observation c7f7b1aa-9db5-41af-b91c-70601291b98c · outbound

This paper cites S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.084316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.621410Z digest=sha256:b6312731ee67956dd2a6d7cb3781943bdfd8fd752e86b0d532684b77420a669c

Observation 367733f4-e4ea-4cee-bbcc-fb3d3563caa0 · outbound

This paper cites and Dobriban, E.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Dobriban, E

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.070732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.624853Z digest=sha256:dfc2d17890556505c919a39056ef6845f9239cb2ad57d23902cdf8cb488d07b5

Observation b12ff897-c3ca-4236-9049-ae560dbd4a27 · outbound

This paper cites E., and Makhzani, A.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning E., and Makhzani, A

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.057485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.628536Z digest=sha256:2cf62c9d76ce169f4ca16ea4ddfa9121485274ea591e5a0db49228f03cb31e92

Observation 388f6a38-5390-42c8-9c6d-1631db713a18 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.632426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.632426Z digest=sha256:572d2d5581e7a4594702a42d64fea73fe7033dcebc6320ced9be8b3204c57f2a

Observation d154584e-0edb-45c6-8104-9eb67556d371 · outbound

This paper cites Deep learning via Hessian -free optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Deep learning via Hessian -free optimization

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.037180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.636203Z digest=sha256:d92ea61cd8a1470a591b06c9ff936f19d8305df00077ed06cc0fa418e994d20a

Observation a8a14744-97ca-4df6-982b-504df2350ef7 · outbound

This paper cites New insights and perspectives on the natural gradient method.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning New insights and perspectives on the natural gradient method

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.640176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.640176Z digest=sha256:90fc0014cf683eb230d5dc3f594d8d994de7fe11597ddd43c0657c9b34432e95

Observation d26dd43c-c00c-44eb-99e2-7541db46b928 · outbound

This paper cites and Grosse, R.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Grosse, R

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.016984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.643973Z digest=sha256:48262db671e2813d694c106d3db4178b966ee9d6720c7035688ed096fdd606a1

Observation ea6facc4-7413-47de-9477-bb75f5272864 · outbound

This paper cites The benefit of multitask representation learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The benefit of multitask representation learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.004539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.647750Z digest=sha256:4f018c1967012cd91347b21fa549416e6717fe201284cc3c2172e9f6069052f7

Observation ca064cd8-24a5-4eff-a8e5-82afb1ec2bed · outbound

This paper cites and Montanari, A.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Montanari, A

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.651377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.651377Z digest=sha256:5ff876bc9312a1f4b1012cc61ddcddefe0788f1b877068b402ba0a5a3a98e94b

Observation 72c85b7c-06ff-4720-8cb5-7bc8c702ff7c · outbound

This paper cites Announcing the results of the inaugural AlgoPerf : Training algorithms benchmark competition, 2024.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Announcing the results of the inaugural AlgoPerf : Training algorithms benchmark competition, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.984276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.655175Z digest=sha256:923789a35fbfb882975aeb4f70478bbdcd6fce28b7a31827478167cca7924e05

Observation dfc3bb58-f64e-4136-8191-d2ceed24bcf8 · outbound

This paper cites Asymptotics of Linear Regression with Linearly Dependent Data.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Asymptotics of Linear Regression with Linearly Dependent Data

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.658970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.658970Z digest=sha256:89045ae00d045b81e29eb879ce92ccf6e73601ad2ec90bfe22efbc606568d2e4

Observation 49b4fddf-ba51-4679-b613-98ca54407e0b · outbound

This paper cites Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.662641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.662641Z digest=sha256:d4e5ba1be7ee340fe72af566c3f8d49302297ce3a7a3f85ffb6d2f1d265e7e34

Observation dc3c028a-32bc-490d-b3cb-245003e03137 · outbound

This paper cites A theory of non-linear feature learning with one gradient step in two-layer neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A theory of non-linear feature learning with one gradient step in two-layer neural networks

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.971707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.666465Z digest=sha256:158375daedebe119f6d816214ae648795db7bde1ff3aa0a3e564912419debc8a

Observation cd2205b1-aa50-48c0-b8f3-5eb254c2b041 · outbound

This paper cites The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.669988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.669988Z digest=sha256:e4a3d8a4daa9d5ad9f368e21037dfc0ea2d70533d6fac35773c1504a618aa9f6

Observation c3a67bcd-ac56-446b-b883-d372f0cb6099 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A New Perspective on Shampoo's Preconditioner

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.673899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.673899Z digest=sha256:f0e2b2b0102be6e61ba40882d3f61cd01fafefba0bfbbf80989e92f804bedc8f

Observation 5c6387ee-d047-4a34-8b20-226e3634757a · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.959520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.678103Z digest=sha256:6768399282b453b82d786991ffaec6bb0e54d75925fb8500f20a85b9b00ba237

Observation cd8868c1-1a99-41ee-9b96-5036ab4b028c · outbound

This paper cites The Effects of Multi-Task Learning on ReLU Neural Network Functions.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The Effects of Multi-Task Learning on ReLU Neural Network Functions

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-09T14:47:42.212659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.681743Z digest=sha256:2ffceece1ad559cdcf32e5b2f8fda67a1ca4b12a6ccdc2dcdd2b950800ac372a

Observation e8b91276-6fb6-4dcf-b223-3a0ebf4c5970 · outbound

This paper cites and Vaswani, N.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Vaswani, N

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.947872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.685732Z digest=sha256:8d4d2ec7784d272a2fee1a1f2c9969fb91597e6e65328ae984a8c79b31a2f1f9

Observation 231319a4-4634-44bb-b4b2-d86129a81411 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.935618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.689591Z digest=sha256:41cf2bdab48cc644eca3b4e1d60ad32b31ca7cde9d8273bda79ba101f9c62ee4

Observation 842238ef-a75e-4a88-8da8-3d2c4826019e · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.922094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.693289Z digest=sha256:221b7dfb4f5e0fdcb5ebc5411c1d54d05c0e9a48254c089afb97300389dc8e69

Observation 59412849-2ae1-4f39-9488-9994150bfc71 · outbound

This paper cites and Wright, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Wright, S

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.696612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.696612Z digest=sha256:bf3d9688eddbb573387b3a2ca93c787320bcd42213cc6fd52ce12c4745de05b7

Observation c8c6eed3-3c23-4b61-98ce-4e57ab071899 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.700012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.700012Z digest=sha256:315f888da2328473d4277a2e5c83018ab67d85b76e9ce7a2b8aab51e81915723

Observation cb2ddf47-a413-4315-88d7-fc3902713327 · outbound

This paper cites and Recht, B.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Recht, B

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.896569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.704030Z digest=sha256:77376d76e08e417e3013452112e12ac9f200a9ebededa866af6231d9d115efea

Observation d88406f9-136f-404b-a8b6-3ccc80639d69 · outbound

This paper cites J., Kale, S., and Kumar, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning J., Kale, S., and Kumar, S

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.884738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.708947Z digest=sha256:61586af193b9e2a385dbfbdfc41b6cd48fc7053b7e7028e9959c2a693aed50f3

Observation af7d01e1-2641-4f1c-a187-18c9bc667570 · outbound

This paper cites and Vershynin, R.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Vershynin, R

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.872327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.712662Z digest=sha256:80f65fe81694258425cf1df53ca548c5915f962d0ad3b5e81766907320ddc291

Observation fd2c4366-ab72-4c83-92f8-fe022a0627a5 · outbound

This paper cites M., Schneider, F., and Hennig, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning M., Schneider, F., and Hennig, P

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.861461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.716403Z digest=sha256:7667c26d09fda4073bf904bd4652b841f7857f70b757024f3bacf6fb8a94e51e

Observation 6e58c8ab-04e5-466d-9273-3cf1e012a7ed · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.850858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.720427Z digest=sha256:de9eb861584739bf25628fe0a9de421f0a064b3437f649da64718f67dba280af

Observation 5ce83e7f-3c9c-4240-9c0a-ed5c29df1bf1 · outbound

This paper cites A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.724257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.724257Z digest=sha256:1ccdfb911257d4a0dadac7a21f7b1c928cd98c808fb09a281cb3386f709be10c

Observation c59315d4-4c51-4e1b-a509-6f4a23834c4f · outbound

This paper cites A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.838365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.728250Z digest=sha256:b63185c5fe32bf344734aab1ab722949a77ad0b7c2c163fe5c8449025d9be23f

Pith citing papers

Observation c594bbe2-d8d7-4cec-bd63-908906f78eb9 · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:26.668562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:49:26.668562Z digest=sha256:74eee6bb9f26ea48f602992ef3175c08f002b3230e0689e25975baec5beee7ed

Observation 5512bc06-639e-46bb-ac2b-6717516fa8eb · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T04:23:02.017903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:23:02.017903Z digest=sha256:b466d84f4488f324248674059121cb2a8b7ffee4bc1a5a94ca40cb3f21ea1b67