Pith. sign in

Paper Citation Record · LEDGER

nGPT: Normalized Transformer with Representation Learning on the Hypersphere

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2410.01131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01131 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:23:40.961082Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:30:07.817010Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 387a7cb2-6a81-4870-b072-6f18762890be · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.961082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.961082Z digest=sha256:8c545f6132a711dc03fac8659aa014623887b8e47e599b057716a515d3a27d2e

Observation 88573fc3-ad48-4eaa-b307-dfd6c924cc3e · inbound

Normalized Matching Transformer cites this paper.

Normalized Matching Transformer nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:05:11.573058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:03:16.161201Z digest=sha256:c92697c59e050d03a559aab2e6168e16ca52ec3a0980e09f75218766fdd0b201

Observation 7b47b369-0bb6-4d15-bd65-c14e732def0f · inbound

Superposition Yields Robust Neural Scaling cites this paper.

Superposition Yields Robust Neural Scaling nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:36:25.534812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T14:38:44.789822Z digest=sha256:cfc5e53d52e75b06c4f9115945b5f3233a5167b91c7d4623716a744ea1eb6a27

Observation 3d9ed69f-484d-48ca-8a46-892eb45a7214 · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:01:46.189690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T18:56:48.722344Z digest=sha256:e154f3f805241b3f67e6f64fc3a800e152f303645e18e82ea47937479e194645

Observation d37b01ab-8058-4db7-8d15-a58eb8b1d551 · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T10:25:17.146093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:25:17.146093Z digest=sha256:614af2597bec9482596270a0e5072c8e8001c318e9379ae973e10f6ac5551e84

Observation 31d74673-2c9c-49e5-bac5-b2bbd047cde4 · inbound

Universal One-third Time Scaling in Learning Peaked Distributions cites this paper.

Universal One-third Time Scaling in Learning Peaked Distributions nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:01:12.799568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:01:12.799568Z digest=sha256:4ac914a4120b8b65195638757fafedea658ade012795e9187f32eeee6f1f3c1c

Observation d37d4404-41ed-4d4c-afe2-b8aaa8b02984 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:49.967839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:b7309f26b3ca3064d5957ff63ae94ebae0a137b25cd79f8ba7e089bd7b3744bd

Observation 1995a4cf-7eb9-4beb-9620-08691b5fd5f0 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:12:41.324396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:45ff2decce881f9987af62c34cac5439d2015a23fb0fd1111973f07446c6fe31

Observation 116c1d4c-9209-4a51-ae74-c921b6830765 · inbound

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer cites this paper.

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.441727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:29:58.495518Z digest=sha256:3ab2985f47233d8c318ce3e8f41e82182ef4cf810a339ae4e197138aeaa08626

Observation e5ba9dbe-d9d5-43d1-b2eb-7cb27ed5d92b · inbound

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning cites this paper.

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.092346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:05:21.069489Z digest=sha256:5f8dd626390b61692f6e5113b8ec22f88d618b28099df966262f1317373950a1

Observation 2f6df2fc-544a-4582-b3a9-75b4789c61fe · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.766353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:df798f790173b716ab38f90f98904899630cedb6d64e7c4e50f543d6cd007360

Observation 740f07ba-42f1-4ad7-97f0-13eb8ed12c3a · inbound

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity cites this paper.

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:08.524404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:11:04.146711Z digest=sha256:974b4f9e7fb6538a28ff1848e2d78c852ca216c8221a6f49e39285a0a7e222f3

Observation b627a3f8-9cea-4adc-9bcf-1d5ab827f2cf · inbound

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives cites this paper.

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:37:19.228142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:34:48.195468Z digest=sha256:7b98be33f35f980f68634c8842339edee5702df7c44a1812af6a897677821867

Observation 7db5a631-d2b0-4884-94ad-538826e3c8e2 · inbound

Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction cites this paper.

Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:57:53.678529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:55:19.362468Z digest=sha256:ce8eaa1de21f4ed0b246fc2db22d2dfb75083bfe7bdd9717dbbdf3e4b38ee6b5

Observation 65f73d7f-438a-4fc4-9137-c19d6dcc5240 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.656032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:a2d91ce8843f60d5333697b0e8cf36f8ef4604fe630ea824b6e7fd6c7482edc4

Observation 1cb962d3-3aad-45f7-a761-3b2b78183328 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.818744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:41cd877399695b5668462b4aa35532522215a9ea3ff077f14995c7205439b126

Observation f2e2411a-d06e-4433-add0-6b0d47495558 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:09.204824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:09.204824Z digest=sha256:6935f1e18bf31bcdcda48a2872f01728fea2f5d8f4ba1fc046af893053e0669d