Pith. sign in

Paper Citation Record · LEDGER

Theory, Analysis, and Best Practices for Sigmoid Self-Attention

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2409.04431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.04431 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:15:39.513319Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:17:24.595834Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c4a8870-83d9-4669-b8e3-f76ba36b05fd · inbound

When Attention Sink Emerges in Language Models: An Empirical View cites this paper.

When Attention Sink Emerges in Language Models: An Empirical View Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:41:03.735502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T17:41:03.674759Z digest=sha256:fdebaffac1f47521a932e37f5e1ae52b06635c3be02b8ca27dd162c4d6ad34ac

Observation 696ec69e-2f2b-4cef-a77d-81ec2fb280fa · inbound

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models cites this paper.

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T16:15:39.513319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:15:39.513319Z digest=sha256:fad9d92cc429805c45a3300c9a8fc154bd978142cf0253515b2f19c45a2b0f5c

Observation 8dc96183-e38f-4ce2-9f52-f5f2eb4a0572 · inbound

Smarter Together: Combining Large Language Models and Small Models for Physiological Signals Visual Inspection cites this paper.

Smarter Together: Combining Large Language Models and Small Models for Physiological Signals Visual Inspection Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:23.796832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:42:23.796832Z digest=sha256:8f9353930e20499ba1ffdedd750ce8f7d1de73efc8d3f617d8433b08a3c724fe

Observation 8d03266f-9e1e-4f01-b669-d63b7e13e42e · inbound

From Molecules to Mixtures: Learning Representations of Olfactory Mixture Similarity using Inductive Biases cites this paper.

From Molecules to Mixtures: Learning Representations of Olfactory Mixture Similarity using Inductive Biases Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T13:41:14.511574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:41:14.511574Z digest=sha256:3e56cce915166b0f3f181d698652405b2cf412b4208f5c9af78d059bb3610fde

Observation 31572126-1d04-47c5-a89a-8b8c22a77a88 · inbound

A Unified Perspective on the Dynamics of Deep Transformers cites this paper.

A Unified Perspective on the Dynamics of Deep Transformers Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T00:04:10.926525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:04:10.926525Z digest=sha256:0f1f83d2925d160219dc3a0bb83e6a6c2c4a3fef5bd59af8d7ac097fa8a63798

Observation 83fe0c0b-b786-43e1-a462-f18671deb5dd · inbound

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective cites this paper.

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:40:10.082924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:40:10.082924Z digest=sha256:46e79b6cccc612fd30cd65a934ca1822c748c7348538ce6aafe3ab9e2d0c86dd

Observation 54681b64-712b-4e5c-83a1-834e2d45ba3d · inbound

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach cites this paper.

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:05:31.676260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:05:31.676260Z digest=sha256:0565cd71d6ccbe5ae3d99db90fa047ac81abc1ee8607c87ba88c9d42741394d2

Observation 1fbe8f8a-bc61-4074-905e-185e3b70a6d6 · inbound

Systematic Outliers in Large Language Models cites this paper.

Systematic Outliers in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:37:37.515845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:37:37.515845Z digest=sha256:487df34d0b05aef491b53f72527c92b89eccca6794279220b2809cf30336c4b6

Observation c125d649-43ce-4369-910f-d547675c7ca4 · inbound

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization cites this paper.

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.381114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.381114Z digest=sha256:2cc508011a57cd2083ad5ce663f24ed1616e2a11ca777daacb3892ba80b2a290

Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.778678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.778678Z digest=sha256:fa7fca67684c72ce76c18e4a564fd75e6f14f0e097d39485230574f11bac0d33

Observation ffc2af75-1f8e-4578-b29a-6f39eeb387ac · inbound

Scaling Context Requires Rethinking Attention cites this paper.

Scaling Context Requires Rethinking Attention Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.782864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.782864Z digest=sha256:4f316b2c4a70cfe1f339221aad42d985d078c12e62701d326ff00d28f2367fee

Observation 3b93fb8e-4c74-4dcb-aca6-d41795f98465 · inbound

Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations cites this paper.

Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:10.426250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:10.426250Z digest=sha256:da37834fba1a45d98480c84702f6b699636a529ccec82fbec4be8988f39c8c6c

Observation e38e9aa0-9fda-4b27-abae-cde295d70edc · inbound

Capacity-Controlled Global Attention for Graph Transformers cites this paper.

Capacity-Controlled Global Attention for Graph Transformers Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.767360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:27:24.557866Z digest=sha256:478479e64e8cb35d00d0e92b43141033aa06aa611ef6d2c6e241034081ff3962

Observation f86867b3-92f9-4a80-a47b-29f35ab9a31c · inbound

Cubit: Token Mixer with Kernel Ridge Regression cites this paper.

Cubit: Token Mixer with Kernel Ridge Regression Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:06:10.787346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:38:19.925573Z digest=sha256:cfcd6ae2f2fb9dcb5a1894a83639958ba39f0db3ec3b1d913795425477f11c1f

Observation b7950fc4-5026-4447-862f-bced4a513992 · inbound

Cubit: Token Mixer with Kernel Ridge Regression cites this paper.

Cubit: Token Mixer with Kernel Ridge Regression Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:11.009258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:34:36.108826Z digest=sha256:04c5b4dd2669705243aaea80282926cf0ac66498f79f104d2d991870784e1191

Observation a8f9fc8d-b472-4959-8fea-6820f08824cb · inbound

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity cites this paper.

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:08.516991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:11:04.146711Z digest=sha256:cd947092d79a36f7acbb871fe5e235baba50a85df3c27a705318dd5e21dac091

Observation b1fd8067-78c2-4245-94e1-f0862b4855c7 · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:40.797845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:29:20.796512Z digest=sha256:2b382032d88d46757b4694abe0c528cbfd591f8cb0be6aaeb909bdd1e63074c1

Observation 2ae133e3-518a-4b55-8959-1c2719b83d1e · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:19:29.111658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:03:25.624300Z digest=sha256:6a8059118a4628b057a5b2f1f4f0d33a5b4561227c5087a56b97fbf3b36513da

Observation d20d30e7-4356-401e-b839-9e29cde7c142 · inbound

Complex-Valued Phase-Coherent Transformer cites this paper.

Complex-Valued Phase-Coherent Transformer Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.147412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:34:20.893124Z digest=sha256:a33ef4dfaa7bc4c5953898eeafa69a1d6b2fa43598d9ff0c9d7baa3b98a15b5d

Observation d057774a-4e16-460f-8f77-6f7cac95eb0f · inbound

Complex-Valued Phase-Coherent Transformer cites this paper.

Complex-Valued Phase-Coherent Transformer Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T14:31:47.302932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:31:47.302932Z digest=sha256:e269033fa2b0f29789af551a08ab4a72219152601685877f8a42e969f8a795c6

Observation 16d1e8f8-de0a-4812-82f2-b7527e32b3e9 · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:42.072047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:0c7061c451adb98bdadfd26ca76780eaff5b8457be8f3e9621d0ee0ac78e6a91

Observation 830e1fcf-1c26-46f4-8683-29111e1fc631 · inbound

Towards Understanding Self-Pretraining for Sequence Classification cites this paper.

Towards Understanding Self-Pretraining for Sequence Classification Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.908338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T05:29:58.809024Z digest=sha256:9e8db73f7372be321337d86186f929481d4b2a529fca100cd741a6634636eb39

Observation 9c37073b-4edb-41c7-8e1f-4e2442124a49 · inbound

Forget Attention: Importance-Aware Attention Is All You Need cites this paper.

Forget Attention: Importance-Aware Attention Is All You Need Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.024853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:38:40.948032Z digest=sha256:6e62b1b65612084481289bc29291eaeab87439c30db10b072696c83579f8a9bc

Observation a2f19bb8-9c0a-4305-87cd-56ab4aabf127 · inbound

A Unifying View of Attention Sinks: Two Algorithms, Two Solutions cites this paper.

A Unifying View of Attention Sinks: Two Algorithms, Two Solutions Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.597286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:51:41.329591Z digest=sha256:8c6d831aa731f6a9deeaca05bffcb092fcbd185d287fd18a2b7fa093aa38c75e

Observation b4ea3e01-b25e-48db-861b-8dc620cab92f · inbound

Legible-by-Construction: Attention and End-to-End Transformers cites this paper.

Legible-by-Construction: Attention and End-to-End Transformers Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T20:06:55.102341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:06:55.102341Z digest=sha256:d6fa91df28346316b007695d4ee9b29ac4bdf4e73a64874a8552927160c93e63