Pith. sign in

Paper Citation Record · LEDGER

Theory, Analysis, and Best Practices for Sigmoid Self-Attention

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2409.04431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.04431 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:04:10.926525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:17:24.595834Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c4a8870-83d9-4669-b8e3-f76ba36b05fd · inbound

When Attention Sink Emerges in Language Models: An Empirical View cites this paper.

When Attention Sink Emerges in Language Models: An Empirical View Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:41:03.735502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T17:41:03.674759Z digest=sha256:0de0ed9974c6b725079a779638b0afe47d5384af6773801c55d3dde0e69c567d

Observation 31572126-1d04-47c5-a89a-8b8c22a77a88 · inbound

A Unified Perspective on the Dynamics of Deep Transformers cites this paper.

A Unified Perspective on the Dynamics of Deep Transformers Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T00:04:10.926525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:04:10.926525Z digest=sha256:3b04434d3ae77d202b8fc3b59ec7cd8889a3fb462d5cab6ede62eec5211c007b

Observation 83fe0c0b-b786-43e1-a462-f18671deb5dd · inbound

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective cites this paper.

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:40:10.082924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:40:10.082924Z digest=sha256:fe74b3375c0a9553e4d9a1b8e111af2fd831e31a8b2124b5c0196177b63ccdbe

Observation 54681b64-712b-4e5c-83a1-834e2d45ba3d · inbound

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach cites this paper.

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:05:31.676260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:05:31.676260Z digest=sha256:7e8cb72beb20a20bc8460f01eae6edbfbed5aaa9fa0a8c1c8ae44f0ee4f691a1

Observation 1fbe8f8a-bc61-4074-905e-185e3b70a6d6 · inbound

Systematic Outliers in Large Language Models cites this paper.

Systematic Outliers in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:37:37.515845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:37:37.515845Z digest=sha256:c9245c0f981333889941ff22a9e35f7ebf77fa48a03941f57c3f3c4b91096e02

Observation c125d649-43ce-4369-910f-d547675c7ca4 · inbound

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization cites this paper.

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.381114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.381114Z digest=sha256:6bf02496605014f499ae0b1457c37ef0a5026ebe8181c694900330ee9907e70c

Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.778678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.778678Z digest=sha256:7bd63a948c4bef39103c3604e111f9469e9a9b70e126b7b8a450a1f3333e873c

Observation ffc2af75-1f8e-4578-b29a-6f39eeb387ac · inbound

Scaling Context Requires Rethinking Attention cites this paper.

Scaling Context Requires Rethinking Attention Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.782864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.782864Z digest=sha256:1698bb2e41f17ed03651b586eea618c546a6e97d806176b06a82a0f6bb02fe75

Observation 3b93fb8e-4c74-4dcb-aca6-d41795f98465 · inbound

Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations cites this paper.

Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:10.426250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:10.426250Z digest=sha256:1efd5d46d697beb40d6ab2593229907a4ed1509443f48fbd6a5efc56cd1ef12f

Observation e38e9aa0-9fda-4b27-abae-cde295d70edc · inbound

Capacity-Controlled Global Attention for Graph Transformers cites this paper.

Capacity-Controlled Global Attention for Graph Transformers Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.767360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:27:24.557866Z digest=sha256:c3fa30bbde8b0eb0d4785e1d7bc302c4416e561b3c1e9c3d096da964bfbb0ea1

Observation f86867b3-92f9-4a80-a47b-29f35ab9a31c · inbound

Cubit: Token Mixer with Kernel Ridge Regression cites this paper.

Cubit: Token Mixer with Kernel Ridge Regression Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:06:10.787346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:38:19.925573Z digest=sha256:159317937fc54cf0c70a9468dfed6bc9c3deeffc33d47701834a44d1f1ae78e6

Observation b7950fc4-5026-4447-862f-bced4a513992 · inbound

Cubit: Token Mixer with Kernel Ridge Regression cites this paper.

Cubit: Token Mixer with Kernel Ridge Regression Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:11.009258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:34:36.108826Z digest=sha256:3a6db31124b0c059248ae9189301ee252eaadc6f7f928bc735905890b6f77137

Observation a8f9fc8d-b472-4959-8fea-6820f08824cb · inbound

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity cites this paper.

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:08.516991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:11:04.146711Z digest=sha256:a358e18bc68d814a53914f4f3039a7d9a35b172ed149cf2024c9c1ab724d7109

Observation b1fd8067-78c2-4245-94e1-f0862b4855c7 · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:40.797845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:29:20.796512Z digest=sha256:f05d4e70736ee216d63a80a25eca97c446f9a5f10752a9addb7c027a94a17fc6

Observation 2ae133e3-518a-4b55-8959-1c2719b83d1e · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:19:29.111658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:03:25.624300Z digest=sha256:052729cde54c4f5682db95e5e467357a1dbb3d1a837a81ebafef2b3f472115dc

Observation d20d30e7-4356-401e-b839-9e29cde7c142 · inbound

Complex-Valued Phase-Coherent Transformer cites this paper.

Complex-Valued Phase-Coherent Transformer Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.147412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:34:20.893124Z digest=sha256:e55e335d4e52e4450f95dcf6cbd899378831907a069d8321f02b850c185ca306

Observation d057774a-4e16-460f-8f77-6f7cac95eb0f · inbound

Complex-Valued Phase-Coherent Transformer cites this paper.

Complex-Valued Phase-Coherent Transformer Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T14:31:47.302932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:31:47.302932Z digest=sha256:7dddcf0e95f2c00498815bec4ed88256ff2e577f358a8203815822e8d57aeab5

Observation 16d1e8f8-de0a-4812-82f2-b7527e32b3e9 · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:42.072047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:d00228bdde1fa6f3f9fe98787922c65a3b66f8cbfd6e250f5d669d4b751aabc1

Observation 830e1fcf-1c26-46f4-8683-29111e1fc631 · inbound

Towards Understanding Self-Pretraining for Sequence Classification cites this paper.

Towards Understanding Self-Pretraining for Sequence Classification Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.908338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T05:29:58.809024Z digest=sha256:4a804667555f56f3922da7e3be71102164969dbbb20da9ca39bd9273fb197c21

Observation 9c37073b-4edb-41c7-8e1f-4e2442124a49 · inbound

Forget Attention: Importance-Aware Attention Is All You Need cites this paper.

Forget Attention: Importance-Aware Attention Is All You Need Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.024853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:38:40.948032Z digest=sha256:07a18cf11920eb7e3cb79e1339b17a53ec5c7648d023550da3e8c5df4b44935e

Observation a2f19bb8-9c0a-4305-87cd-56ab4aabf127 · inbound

A Unifying View of Attention Sinks: Two Algorithms, Two Solutions cites this paper.

A Unifying View of Attention Sinks: Two Algorithms, Two Solutions Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.597286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:51:41.329591Z digest=sha256:ae00b3a1504af4d1d4d9483edf89eff8b690951046588d041da31803319849e2

Observation b4ea3e01-b25e-48db-861b-8dc620cab92f · inbound

Legible-by-Construction: Attention and End-to-End Transformers cites this paper.

Legible-by-Construction: Attention and End-to-End Transformers Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T20:06:55.102341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:06:55.102341Z digest=sha256:a1bd16d70b38fb84f3201e6d58b9e384892026533b4d09fe076810682183cf56