Pith. sign in

Paper Citation Record · LEDGER

How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2305.18270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.18270 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:32:47.947181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:44.997380Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 30473510-18fe-4bff-b315-a1aa1c0189ab · inbound

Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks cites this paper.

Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:47.947181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:32:47.947181Z digest=sha256:e3a2a6fe41c7736ef46ddc86d61ae108b8b70ebfe021156bf1fcefd0bb907bad

Observation fa86431a-df41-4314-9fb5-a8535fcabfaf · inbound

Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories cites this paper.

Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T04:36:32.044680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:36:32.044680Z digest=sha256:a3081f6e5954b3b6ad4880e3eaf3eed2dace55fb2535c4e67b03fe49d3b55909

Observation a261ff52-bf6c-4edb-a572-9f73bdb300c6 · inbound

Low-dimensional Functions are Efficiently Learnable under Randomly Biased Distributions cites this paper.

Low-dimensional Functions are Efficiently Learnable under Randomly Biased Distributions How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:34:16.642638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:34:16.642638Z digest=sha256:083de74230d6ffdefed786ae66a83259ac8e6d65db3b7e40c5ea61db56027f1c

Observation 538f54cd-b6d1-4db0-a809-544ffa8652d4 · inbound

On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective cites this paper.

On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:54.621375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:54.621375Z digest=sha256:e5bbe825ad8580ca11ec17e2eb482991b351ad7554cc0bbcf264fa01174daa74

Observation b92bcd50-dfe8-41fa-8cef-b6b523db09c1 · inbound

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability cites this paper.

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:20:52.659015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T11:20:33.400885Z digest=sha256:78a20ac3b85778b0d046eba186d775dc2e7157d794edfc0373884d4a348c8e5e

Observation eec2132c-6579-4d21-98d0-7bf1a82f197e · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.999214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:82a6ae1e8a508d3bc6872135995666e8b8993bbbccf0c4dda876c9a740e3aa84

Observation faf00f46-23d0-4e6b-a86c-1129ebdfc471 · inbound

Spectral phase transitions and trainability in neural network learning dynamics cites this paper.

Spectral phase transitions and trainability in neural network learning dynamics How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:47.552849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T01:22:17.359656Z digest=sha256:64cc9b8c67aa30c79bb5e42409a480de24d65d256aa739ea268e919770783310

Observation 35c1a2cd-d361-4f9c-bc7f-b499cafbcb1c · inbound

The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning cites this paper.

The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning How Two-Layer Neural Networks Learn, One (Giant) Step at a Time

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T03:07:54.001891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T03:07:54.001891Z digest=sha256:baed941013d756b9e7328c62eb5bd19a095fb1b3f60c0514529956d2ef61410d