Pith. sign in

Paper Citation Record · LEDGER

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

As of 20 August 2026, this Paper Citation Record lists 100 of 124 outbound references and 2 inbound Pith citation observations for arXiv:2502.01763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01763 v1

Coverage vector

measured 100 of 124 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:40.728250Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:23:02.017903Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 124 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bce02749-8550-4b65-9051-fb87ac1530e8 · outbound

This paper cites and Szepesv \'a ri, C.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Szepesv \'a ri, C

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.304467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.304467Z digest=sha256:8a42d16b9fa6587c9f5444d2321ee51043fdfc3fcf813be6220495eea97f6df2

Observation 7193376e-7dd6-48c3-b2a8-b33e0f8e89e2 · outbound

This paper cites B., and Misiakiewicz, T.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning B., and Misiakiewicz, T

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.309277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.309277Z digest=sha256:e335b2b17f6e7709e5f9f0b69b8af7cd69177f31fa8208a12a6c6bcf1e575573

Observation 2f320a98-16b5-4740-b412-3c65974b592d · outbound

This paper cites B., and Misiakiewicz, T.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning B., and Misiakiewicz, T

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.313480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.313480Z digest=sha256:49e6199ec9e8b80986c766dea02000e6b4334e191063238a00fca5fe991df88f

Observation 1f82431a-f5ef-45a0-919a-6a9adada6074 · outbound

This paper cites Optimization algorithms on matrix manifolds.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Optimization algorithms on matrix manifolds

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.317319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.317319Z digest=sha256:c98961810f94b666fdb91e2431498391042389dc9c2a853d00ad07e4fdc3831c

Observation de925ca4-f8c1-4d47-bbbd-fa25391ce57f · outbound

This paper cites and Pennington, J.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Pennington, J

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.321442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.321442Z digest=sha256:22d04918313770df69dec210ee0c8f55ed6d18be7d78e9bca691aabd644db82b

Observation fbb44cb4-3cd9-4723-9b84-6abc71b8a4a4 · outbound

This paper cites and Parrilo, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Parrilo, P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.325138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.325138Z digest=sha256:36541bf184f3a25b78cf63b16aafad519a483013dd44783bb3b4e75927b9b813

Observation bc46bb90-4b9e-4c1f-8db1-4c69c55aa12b · outbound

This paper cites When Does Preconditioning Help or Hurt Generalization?.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning When Does Preconditioning Help or Hurt Generalization?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.329559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.329559Z digest=sha256:6ea591493e38657e830805a50970bec9d6333ac289ba337c251ba1f32f395005

Observation 4d693f72-8cd8-44b1-b08b-cd2ea91164e3 · outbound

This paper cites Locoprop: Enhancing backprop via local loss optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Locoprop: Enhancing backprop via local loss optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.334202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.334202Z digest=sha256:1df6ef1769b472e067ea612165b23ed8f0ced4d8e05ea6ce98953d09a3de0503

Observation f6ee63cb-87d5-46c8-b7e5-4bfbe809579d · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Scalable Second Order Optimization for Deep Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.338674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.338674Z digest=sha256:686f7be2da341e38150104e2cbebff81f9c28ebbd4dc0b6d203f4202903caa47

Observation 10a1e903-c85c-402b-b5c9-ed75f1db0ffd · outbound

This paper cites Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.343242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.343242Z digest=sha256:d104e22a80c8ee1ba631493996758a3eb21aec60ebe64b27a27500012a7c96e3

Observation ffc32a11-37db-4fb2-84a0-6b7db5fd17ad · outbound

This paper cites Implicit regularization in deep matrix factorization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Implicit regularization in deep matrix factorization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.348170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.348170Z digest=sha256:7da647d652ed9eac991794d0e2de15c0272233f2b1d5358346d922710e3f1a29

Observation 9e0fea14-736a-46a7-85e0-4113327f6a78 · outbound

This paper cites B., and Martens, J.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning B., and Martens, J

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.354122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.354122Z digest=sha256:75fcefd1e6d84732cc210efb8814b3a88860854eef3d5decc28c383361f980ca

Observation b3f1fec8-b5c5-4e6a-bedc-373c61660763 · outbound

This paper cites A., Suzuki, T., Wang, Z., Wu, D., and Yang, G.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A., Suzuki, T., Wang, Z., Wu, D., and Yang, G

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.361665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.361665Z digest=sha256:a6b8a8fee555847e258b9ba25b1ecebb976457e9ee45bf9c9fd4f041509328b7

Observation b5fb6e65-fc38-4903-9e15-552dc4f0c857 · outbound

This paper cites A., Suzuki, T., Wang, Z., and Wu, D.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A., Suzuki, T., Wang, Z., and Wu, D

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.367830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.367830Z digest=sha256:02d9dee65fe6fedea298a2edd2e2c8339b02b28a0fd6ba22e881ab1f7b31a6d6

Observation 07808b72-662a-4e0d-aa38-38a34939e8f9 · outbound

This paper cites Scaling laws of optimization, 2024.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Scaling laws of optimization, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.372639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.372639Z digest=sha256:011444fbf375627312ef5e5292b86bd5d21854caab201b8b3b8a73298b523d57

Observation 872b33da-ce0b-410a-b93f-b550286e9611 · outbound

This paper cites and Lee, J.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Lee, J

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.377075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.377075Z digest=sha256:d7527b05a45ea86c88c8bfd966e1fe874c0d0f0237e6d9e6de67fac1c77712ea

Observation 67f30173-e43f-4a91-89fa-d175b3d5ef5c · outbound

This paper cites Hidden progress in deep learning: SGD learns parities near the computational limit.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Hidden progress in deep learning: SGD learns parities near the computational limit

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.381857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.381857Z digest=sha256:5f9c9798cc135fc5250df43a79681c22b46a1f7e76a6285002af18edcedecd4f

Observation 298c1374-61c3-40f4-a042-331e08af5ff4 · outbound

This paper cites Online stochastic gradient descent on non-convex losses from high-dimensional inference.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Online stochastic gradient descent on non-convex losses from high-dimensional inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.386686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.386686Z digest=sha256:e83bb859921e781698d4dda10fd2cba2571ed15aa55a97361c56e217b57ba699

Observation cfa5e545-948c-440d-b037-85207d68765e · outbound

This paper cites Gradient descent on neurons and its link to approximate second-order optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Gradient descent on neurons and its link to approximate second-order optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.391577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.391577Z digest=sha256:7e243bf51c3f207d20af6112880be245c391b55ad518fb3676cbe8b1f69e9899

Observation e104b005-4fdb-4942-9cd6-7ff6a9ed5fc9 · outbound

This paper cites Modular Duality in Deep Learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Modular Duality in Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.399031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.399031Z digest=sha256:50fe2ee9987384d67716907a36cc123154799693c43e8f3b2fff2ae52988b881

Observation b146c1de-38bf-40c4-a319-a6fd06b40fd3 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Old Optimizer, New Norm: An Anthology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.403483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.403483Z digest=sha256:ea167741ea011d991077dd543734fb3985f7df5698b66d278de6acda9cfd8f16

Observation a64b08e4-4b17-473d-8c9e-329c965a33c7 · outbound

This paper cites Learning time-scales in two-layers neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Learning time-scales in two-layers neural networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.408891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.408891Z digest=sha256:4afab115014414b32a1c969435b373c4dac361c98cbcfbf452cd528906e3ae12

Observation 1fbfaf72-6a26-49b0-88a2-c72effb17601 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.412814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.412814Z digest=sha256:30545593e01002d03a52d78ea253eb1b9e25fead9ee616b7c6fbfe0f27d931a9

Observation 30f40d95-194f-4b0e-a58f-bead24c7b0b2 · outbound

This paper cites and Mondelli, M.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Mondelli, M

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.418425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.418425Z digest=sha256:1a6b6bcdb920304d9b2bb832b0141fdfb68e0bce8567915ebb95b8f443810278

Observation bf328dcd-c6a2-4e14-8f98-9b2938b8b042 · outbound

This paper cites Privacy for Free in the Overparameterized Regime.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Privacy for Free in the Overparameterized Regime

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-09T14:47:42.500114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.422681Z digest=sha256:6c0a8da5ee8a8b253687807fc85c7c5fb8756572d3963abfc5845574877dbb94

Observation 517db258-d052-4692-875b-96881237575e · outbound

This paper cites Beyond the universal law of robustness: Sharper laws for random features and neural tangent kernels.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Beyond the universal law of robustness: Sharper laws for random features and neural tangent kernels

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.428124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.428124Z digest=sha256:8934858c430de531fe089496c531097a0e92c374e00e8efe9abcab4078278053

Observation 8a41c11c-128f-4786-90a2-a04690a6f33b · outbound

This paper cites Practical Gauss- Netwon optimisation for deep learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Practical Gauss- Netwon optimisation for deep learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.432156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.432156Z digest=sha256:d1ce5c280fa425e7aa1ed20270d520aee082a677de7aaf0173cce1c9634ea775

Observation d8ade78a-5882-4f91-9a98-047061157a6e · outbound

This paper cites H., Hansen, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning H., Hansen, S

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.435829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.435829Z digest=sha256:27bee5d420c6a1f93e3351045687a948dc716884d5a0b7a695f9ee28d0ea109b

Observation e9f04798-4a48-4bab-a702-aa64104d58f7 · outbound

This paper cites Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.439889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.439889Z digest=sha256:375069e940323aa0772345b9425f875f87d5b0276e84935c5b0b245c4fb001cd

Observation dcdce652-6755-4296-acbc-cf9716c7fb7f · outbound

This paper cites Exploiting shared representations for personalized federated learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Exploiting shared representations for personalized federated learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.443948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.443948Z digest=sha256:9b68666aff8c50147e1c1d1bbca966536a2f845b9db217ed63e277369a845516

Observation 5d5657af-e725-4a2b-a3a8-09f1b66ba8d6 · outbound

This paper cites Provable multi-task representation learning by two-layer ReLU neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Provable multi-task representation learning by two-layer ReLU neural networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.449327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.449327Z digest=sha256:a0645fac93fe14733b149d8f79ec4d092a1edde3b015967a7cbab63fa33fa28d

Observation 256b669c-b153-4ee5-b57c-08d152b1bdbc · outbound

This paper cites Asymptotics of feature learning in two-layer networks after one gradient-step.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Asymptotics of feature learning in two-layer networks after one gradient-step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.453841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.453841Z digest=sha256:5fb6f69ff4d49f92e48422b6154df1961b7c35438c2e6b900ae626cac495258e

Observation fa88d2c3-f48e-4033-8299-440b908b08ec · outbound

This paper cites Benchmarking Neural Network Training Algorithms.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Benchmarking Neural Network Training Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.457590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.457590Z digest=sha256:4a4e60994794c719ac2c7fc9021b5f273b0275aeb5775b6093108b084da940fa

Observation b1bbd743-58fa-4b98-93d6-9628fe76a564 · outbound

This paper cites Neural networks can learn representations with gradient descent.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Neural networks can learn representations with gradient descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.462358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.462358Z digest=sha256:6dabe4c8500c3b46e143f225d03a505172a7309fe27ee254f841915181f732cf

Observation 22cf088e-f81b-41dd-8090-046e598c859a · outbound

This paper cites How two-layer neural networks learn, one (giant) step at a time.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning How two-layer neural networks learn, one (giant) step at a time

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.466603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.466603Z digest=sha256:c94dc10ae31ddd9866c9e127f1c33f55c3c3d71f6070ce7f0cf0be286faa1406

Observation 84927feb-beac-4eb3-a64a-d1fe23e2c2bf · outbound

This paper cites A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.471320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.471320Z digest=sha256:a7dbdf09698d339bb0388a67339c5769ca093ccb1a3d7aad7097ed7f5423f592

Observation 7e01e8de-9c7e-4716-afc4-6d14356ae73b · outbound

This paper cites The benefits of reusing batches for gradient descent in two-layer networks: Breaking the curse of information and leap exponents.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The benefits of reusing batches for gradient descent in two-layer networks: Breaking the curse of information and leap exponents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.475810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.475810Z digest=sha256:040ac06f01cd8cc7fd28ed5c5438c4a054bcaedc90ca4d2a1e84a60392cd1dd1

Observation 73454c87-a7f4-4455-9fd0-ed35153d6aec · outbound

This paper cites Modular block-diagonal curvature approximations for feedforward architectures.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Modular block-diagonal curvature approximations for feedforward architectures

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.480451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.480451Z digest=sha256:571ee14db82b9ee793d2e68c98e5876eb184024cee0e9f53e4cf08424b1b0df1

Observation 6866259e-82f7-440b-84eb-964097cfc13f · outbound

This paper cites Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.486358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.486358Z digest=sha256:da9ad4661be373022e744ce635f8cf9f38507f008ec84b62006a3ddb8bbae260

Observation c762e9fa-410a-40ba-8934-edbd513acfc9 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.490862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.490862Z digest=sha256:1ee51ad87fe9485fa74a85187067da0f16b8616af8271f943cfb8a51057a4c77

Observation 25d7dbda-a607-4810-85b7-8cb65679855e · outbound

This paper cites and Wager, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Wager, S

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.495605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.495605Z digest=sha256:432a379eaa988261ccf44bb4a0b6e2f2ed45db8fb4c52e05de23904757f19334

Observation 15c500be-be2a-4cc5-a5e2-fa8686a15c2a · outbound

This paper cites S., Hu, W., Kakade, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning S., Hu, W., Kakade, S

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.500141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.500141Z digest=sha256:f7be7286de6ba18047fda7a77f2ccd1d51d34b1ed10aa83dc13211ac1d4b0794

Observation bc16ca31-d1cc-4744-85a8-393eb048238e · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Adaptive subgradient methods for online learning and stochastic optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.505091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.505091Z digest=sha256:51d094826ceb848bb44ecf879f6bbcd8ac587d78580ac17fc7ad811721fc63bd

Observation ff4ce0d6-8580-4bcf-8fab-76c3c9ef8204 · outbound

This paper cites Proximal backpropagation.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Proximal backpropagation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.509479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.509479Z digest=sha256:121ee30e3af60b0065df0258c404c8c3c61ba002a8eaf24c85428d3d32192beb

Observation c99f2b0f-2772-4e15-9be3-6dcaa37c2bb9 · outbound

This paper cites Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T14:47:42.442595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.513091Z digest=sha256:459a4a58b68016317cfb39d0f0474c3a23944fb5715ae4cc54603e38c8a86a50

Observation 252bdf19-c590-4bbc-9cd7-29e0d50eb381 · outbound

This paper cites Linearized two-layers neural networks in high dimension.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Linearized two-layers neural networks in high dimension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.516867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.516867Z digest=sha256:bbea68a9c22870294e4130a45e41552c946eba1f852a43a33ef5720e1db261f2

Observation f7931c78-ba89-4082-85ee-c5923442b1b1 · outbound

This paper cites When do neural networks outperform kernel methods? Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12), 2021 b.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning When do neural networks outperform kernel methods? Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12), 2021 b

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.520374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.520374Z digest=sha256:efcf433dc1d3ed9c3b82750fae19618663e439cb288824d8492883a11facbcd0

Observation 85e99ff8-1046-4e3e-a3f8-18dd3eafd38e · outbound

This paper cites A family of variable-metric methods derived by variational means.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A family of variable-metric methods derived by variational means

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.524058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.524058Z digest=sha256:d224d3cac07cfdb000a6fe5364a53304b2df0cd15d3de9a9b5cdea4140b1ea5d

Observation 82c83bb6-3165-4b28-9f16-a163af501fd2 · outbound

This paper cites Practical quasi- Netwon methods for training deep neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Practical quasi- Netwon methods for training deep neural networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.528448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.528448Z digest=sha256:12ec6cf742977dbd989f297e956385d096012e7bf44cafeb43c62cb179e5af4b

Observation abe35d8f-a5d0-4268-8cc0-5d98c753aa40 · outbound

This paper cites The G aussian equivalence of generative models for learning with shallow neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The G aussian equivalence of generative models for learning with shallow neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.533185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.533185Z digest=sha256:cfb452cdfb547219e13bf63a1888e0fc45f87ce0faa64ff0d653602bda0df96b

Observation e457abdd-3640-4945-8bb9-6fc4db3ca515 · outbound

This paper cites Spectral Phase Transitions in Non-Linear Wigner Spiked Models.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Spectral Phase Transitions in Non-Linear Wigner Spiked Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.537284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.537284Z digest=sha256:a9a6ef479756a1e48ebf84ec7e212cf58b67f87617ffb3b8ccf3263b56882941

Observation 421b455b-ee73-471f-a4a2-ca042e11dd29 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.541398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.541398Z digest=sha256:2bec95c2b345a6ed5c3ab12c3dd5271693c969409bc348de9d98d4a336ddcd9e

Observation 8126d448-d040-4a49-b6ac-1eff22992fcd · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Shampoo: Preconditioned stochastic tensor optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.545269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.545269Z digest=sha256:2c1cc7bc8c09ebf771022fcbb90f6b62ba39582b4ca4f6d0e22be91e7b86a36b

Observation bcebb231-9ae7-4250-bd93-c2e08d1f1d6d · outbound

This paper cites and Nica, M.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Nica, M

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.548962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.548962Z digest=sha256:31253bccd0c9f030ae0d6c1add6b5d37c02abda42a81f3a908297d4d09c9c662

Observation 5a6fb108-5b8d-4047-90f3-6f33e8decf73 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.552782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.552782Z digest=sha256:818cd389a0831bdfda52ed0c8a25d04db5caf5454e74ac3eff50fe612c1bd8ee

Observation f2471c44-e407-4422-923a-84655537715f · outbound

This paper cites and Javanmard, A.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Javanmard, A

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.556647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.556647Z digest=sha256:24df2afc5e780c69fb0b5d33afd27a343d15fc0b87fca62a7bb3cbcae96c4ecc

Observation 29217597-fb3d-4d51-ae01-90ea07e95239 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.560608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.560608Z digest=sha256:774b324c0103d792e4081f7c426c4df0666d2e384bee1991f8830adc7b228ba8

Observation 9c12a425-d4ce-4403-aaf6-f07e0d3354ab · outbound

This paper cites M., and Zhang, T.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning M., and Zhang, T

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.215635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.564868Z digest=sha256:821d6fc3a17ea45c7ea6a4d2a3af413917e34a81663ba9516fd7e52b8ecbdb05

Observation bc7af867-b87f-45f9-adec-00bdf6d797ae · outbound

This paper cites and Lu, Y.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Lu, Y

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.201443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.568825Z digest=sha256:90eee15607071eecac1190a12321fe111d0da3bee4d508c98df50646add0300b

Observation f54e91de-1899-4096-adac-1fae999788d5 · outbound

This paper cites and Szegedy, C.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Szegedy, C

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.188677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.572424Z digest=sha256:13b98023b52840fb07b47e2d84a907a461faeb0c37f0298d397c6367db0949e9

Observation f9d9e5e1-2cef-4e89-836f-81fc90171b64 · outbound

This paper cites On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.576362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.576362Z digest=sha256:daa41ca2cbaa8a2cd1c64537a7a0edcda321c5bb441f1c45c1e50907b3d9422d

Observation 690ad69d-fb90-4bbb-b54e-8173f53389ed · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Neural tangent kernel: Convergence and generalization in neural networks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.176504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.580605Z digest=sha256:14e8ebf58f41a6eff08be300ece55f49ec30d755cc2e9b6e44d3f418246ef6bc

Observation 1279429c-1794-4140-be6d-178de71424b7 · outbound

This paper cites Low-rank matrix completion using alternating minimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Low-rank matrix completion using alternating minimization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.163081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.584397Z digest=sha256:d7228a49252bd55fa8825ee27a264f7fc2194e0362145918a95142f5f17fb8c0

Observation 026a4f86-c81a-4f80-9bd7-24b01f956401 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Muon: An optimizer for hidden layers in neural networks, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.147831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.589217Z digest=sha256:2ee24e18528b2ff95b7b15e8aec7144d36599c4ebdebf1f9ccf3937c965d16c9

Observation e77cab59-85e7-4356-b651-1126434c3d34 · outbound

This paper cites Improving Generalization Performance by Switching from Adam to SGD.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Improving Generalization Performance by Switching from Adam to SGD

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.593164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.593164Z digest=sha256:6b186b343619e313c46e4bf290abf34c9e56d99b92770c225484e6bd3b66478b

Observation 652e5d86-f409-41c7-ae50-3cc85dabc55a · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.597248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.597248Z digest=sha256:cca79898c431a0a0d01d0a19cc76d0383db30e68cb05695f82d430c9ac524940

Observation 869a9e7f-5c29-4588-8006-ceb7733b6b02 · outbound

This paper cites M., Ma, T., and Liang, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning M., Ma, T., and Liang, P

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.123219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.601177Z digest=sha256:714353d4c4031ed2c04b135ded95c815c5f29b69904e77fecebe08bb947cef65

Observation 72e137fb-aec8-4d12-af33-abbead0af4b2 · outbound

This paper cites Scalable Optimization in the Modular Norm.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Scalable Optimization in the Modular Norm

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.604767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.604767Z digest=sha256:63aae5a729c8cfc24f92478361616ab44d12f6fa6669b140cf34e2a877d8de9e

Observation 836c0b76-eeb2-4ac8-8076-2ae655001634 · outbound

This paper cites and Massart, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Massart, P

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.109668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.609198Z digest=sha256:f8ca830185d6b8a5d621751af5a38f340a375d060319d52b011b046eb0294560

Observation f975a99a-dee8-4548-9e66-394fcce7a653 · outbound

This paper cites Demystifying disagreement-on-the-line in high dimensions.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Demystifying disagreement-on-the-line in high dimensions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.096942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.613267Z digest=sha256:16ea02a5f426bdafa6fc2b475fe90442e48d61858c3cd77d0d06f2539f255b6e

Observation 4a46ccb1-c848-4c3b-ae50-394da2df9736 · outbound

This paper cites Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.617364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.617364Z digest=sha256:fd2ff2cdf9ed998ea555e0f056f410ff54628620ba241476e98f013852bcbd2f

Observation c7f7b1aa-9db5-41af-b91c-70601291b98c · outbound

This paper cites S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.084316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.621410Z digest=sha256:7459923a2a901e83a1ee320ca5b4131507e8a2e83491bc40694c2aef2be4ebcb

Observation 367733f4-e4ea-4cee-bbcc-fb3d3563caa0 · outbound

This paper cites and Dobriban, E.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Dobriban, E

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.070732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.624853Z digest=sha256:2911c611b5c593bcaae0dae58e0c20c12b52a326b97e2075bc5b072df2506ca3

Observation b12ff897-c3ca-4236-9049-ae560dbd4a27 · outbound

This paper cites E., and Makhzani, A.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning E., and Makhzani, A

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.057485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.628536Z digest=sha256:87e39bec5b92bd2e23b13da7a08411f1acde524b7c10d81edb280dfcde8333df

Observation 388f6a38-5390-42c8-9c6d-1631db713a18 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.632426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.632426Z digest=sha256:dfca24ee6896149971c09fa8374833cff213a831b062dd6517b11243cafa857e

Observation d154584e-0edb-45c6-8104-9eb67556d371 · outbound

This paper cites Deep learning via Hessian -free optimization.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Deep learning via Hessian -free optimization

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.037180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.636203Z digest=sha256:988bc1c2ee3cffb93f130628e0d692a650ba3a4fbdd5692069df6b77bfde20e5

Observation a8a14744-97ca-4df6-982b-504df2350ef7 · outbound

This paper cites New insights and perspectives on the natural gradient method.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning New insights and perspectives on the natural gradient method

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.640176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.640176Z digest=sha256:d136d9a4cd07cb70246c15bd6f8fa3a32c39b56b1f52d3389e716493211c0f03

Observation d26dd43c-c00c-44eb-99e2-7541db46b928 · outbound

This paper cites and Grosse, R.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Grosse, R

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.016984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.643973Z digest=sha256:7b229ba149f26b57bc31c29364ac8efe793b29a40f7f5ade2e2d5d15fe26aa7b

Observation ea6facc4-7413-47de-9477-bb75f5272864 · outbound

This paper cites The benefit of multitask representation learning.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The benefit of multitask representation learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:43.004539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.647750Z digest=sha256:d5cc361830f1eab112e9ec3eb3bb0cb447fb8e599b63c6c8b993d620a3bd1df4

Observation ca064cd8-24a5-4eff-a8e5-82afb1ec2bed · outbound

This paper cites and Montanari, A.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Montanari, A

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.651377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.651377Z digest=sha256:2510940a3bfd5001ef852d374aa493b4919d87e17155f2899ffcc151963ebd82

Observation 72c85b7c-06ff-4720-8cb5-7bc8c702ff7c · outbound

This paper cites Announcing the results of the inaugural AlgoPerf : Training algorithms benchmark competition, 2024.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Announcing the results of the inaugural AlgoPerf : Training algorithms benchmark competition, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.984276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.655175Z digest=sha256:4a9394f698bdc5cef7b46cc3f5ff85836a23607c144553f1b550f0f32acaca55

Observation dfc3bb58-f64e-4136-8191-d2ceed24bcf8 · outbound

This paper cites Asymptotics of Linear Regression with Linearly Dependent Data.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Asymptotics of Linear Regression with Linearly Dependent Data

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.658970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.658970Z digest=sha256:c0cca3a96a0c303cfc6890da5eea8b2ac93d145d6d8d82099aece751fcfc65c6

Observation 49b4fddf-ba51-4679-b613-98ca54407e0b · outbound

This paper cites Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.662641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.662641Z digest=sha256:abe4f786964ea59068d6e7be39bad0f0480e2e2624d50b0c34404ec4ac548855

Observation dc3c028a-32bc-490d-b3cb-245003e03137 · outbound

This paper cites A theory of non-linear feature learning with one gradient step in two-layer neural networks.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A theory of non-linear feature learning with one gradient step in two-layer neural networks

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.971707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.666465Z digest=sha256:9ffd522fc835c04f290c2289ee91ed3cac1c37e3f60739abf360ec7649723c47

Observation cd2205b1-aa50-48c0-b8f3-5eb254c2b041 · outbound

This paper cites The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.669988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.669988Z digest=sha256:6843a9dc8625cb97cca753709406096e3c05b4ce2fdf4110ac586b8c58523e22

Observation c3a67bcd-ac56-446b-b883-d372f0cb6099 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A New Perspective on Shampoo's Preconditioner

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.673899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.673899Z digest=sha256:75a63e3be3f0cefdb8fd1de340a53aabd324666a39c61f1a83ba0f55829fbc8f

Observation 5c6387ee-d047-4a34-8b20-226e3634757a · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.959520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.678103Z digest=sha256:17b508c88077b310cb40bf039dd646843d414e69d89b62b41d86ee96377ef1e3

Observation cd8868c1-1a99-41ee-9b96-5036ab4b028c · outbound

This paper cites The Effects of Multi-Task Learning on ReLU Neural Network Functions.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning The Effects of Multi-Task Learning on ReLU Neural Network Functions

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-09T14:47:42.212659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.681743Z digest=sha256:b906b841826649efff160fab8317374a34b426e94b0d796ccb57a43424fb562a

Observation e8b91276-6fb6-4dcf-b223-3a0ebf4c5970 · outbound

This paper cites and Vaswani, N.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Vaswani, N

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.947872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.685732Z digest=sha256:b03e973a79d8eca396345ffc34fee4c31cfb613e1c4ed8b538b88f5413846ce0

Observation 231319a4-4634-44bb-b4b2-d86129a81411 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.935618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.689591Z digest=sha256:573c56a801bb7fb360d07e293abbf9f5a3f49325472c8b55fdc02da7985d6705

Observation 842238ef-a75e-4a88-8da8-3d2c4826019e · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.922094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.693289Z digest=sha256:4fa4c1d82892815711a3a378534e1eb6d5f0de0516969c3ab20388482ba37db4

Observation 59412849-2ae1-4f39-9488-9994150bfc71 · outbound

This paper cites and Wright, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Wright, S

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.696612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.696612Z digest=sha256:8f26cd45f6b355b055b32980242bdd1ee33e5cb93c2a425e4d69987afca2fa62

Observation c8c6eed3-3c23-4b61-98ce-4e57ab071899 · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.700012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.700012Z digest=sha256:0e4536da326d3ab28905f2d33bee039af752306fbeb8bdb4cf88d45dd6d3cea6

Observation cb2ddf47-a413-4315-88d7-fc3902713327 · outbound

This paper cites and Recht, B.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Recht, B

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.896569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.704030Z digest=sha256:23ff2316b8a433ab7e6fb35f938fdf46a7e552f6f13f38a602fbd7498e42627b

Observation d88406f9-136f-404b-a8b6-3ccc80639d69 · outbound

This paper cites J., Kale, S., and Kumar, S.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning J., Kale, S., and Kumar, S

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.884738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.708947Z digest=sha256:841eb8ff3bda0aacc9c1635eab2306837dff241eda15d0cd3ac59fefffc7b1ea

Observation af7d01e1-2641-4f1c-a187-18c9bc667570 · outbound

This paper cites and Vershynin, R.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning and Vershynin, R

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.872327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.712662Z digest=sha256:231141c63fb06ee069e6d69aedac971402d4391f5b177a98eb8084977a141df4

Observation fd2c4366-ab72-4c83-92f8-fe022a0627a5 · outbound

This paper cites M., Schneider, F., and Hennig, P.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning M., Schneider, F., and Hennig, P

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.861461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.716403Z digest=sha256:17530fda3ba4d72f173b19a3bb3c5c6a447f24f7e45dc44a603e84923d14a837

Observation 6e58c8ab-04e5-466d-9273-3cf1e012a7ed · outbound

This paper cites an unresolved cited work.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:42.850858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.720427Z digest=sha256:0b0bd472a647c214962e1be98a2b4bf712f2994badfed664ea2b0e92493464d0

Observation 5ce83e7f-3c9c-4240-9c0a-ed5c29df1bf1 · outbound

This paper cites A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.724257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.724257Z digest=sha256:2b1a821a7e94aec6847c0acb34d2158c02f53b81c78d6be7e140405cd64c4037

Observation c59315d4-4c51-4e1b-a509-6f4a23834c4f · outbound

This paper cites A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:42.838365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:47:40.728250Z digest=sha256:4246af6d6e78eb7fd8bfa718d733579bed4f23e0ff487ae62cdea6b3bc1c0f00

Pith citing papers

Observation c594bbe2-d8d7-4cec-bd63-908906f78eb9 · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:26.668562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:49:26.668562Z digest=sha256:ccda911432a2ef62c0ab6e998b2fce5b740a51a6e846aaee835d0de2e4c92062

Observation 5512bc06-639e-46bb-ac2b-6717516fa8eb · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T04:23:02.017903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:23:02.017903Z digest=sha256:5c7e56805f321b977c51c36b8c8fc3568d30b858fa3c0bedae58d96a7c6b2ca3