Pith. sign in

Paper Citation Record · LEDGER

Does Self-Attention Need Separate Weights in Transformers?

As of 12 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2412.00359.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00359 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:33:45.978048Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbcbe6c8-9a67-4f24-9517-9b468beac773 · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:33:46.708685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T05:33:45.665416Z digest=sha256:6ddcd146d41fbec3fc0e49ad4ccae83dabf46fcd1d5821d79b9c2f2f63d5ea8b

Observation 556cc437-843c-49cf-be2c-e31782463e45 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Does Self-Attention Need Separate Weights in Transformers? Neural Machine Translation by Jointly Learning to Align and Translate

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.678585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.678585Z digest=sha256:ed686fb88a7c2d64f6dba494ed7cf06060300f025237d771c7e69b0ca947ea1b

Observation c5abf837-af36-4ddf-97f8-af1e1730064a · outbound

This paper cites Longformer: The Long-Document Transformer.

Does Self-Attention Need Separate Weights in Transformers? Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.690320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.690320Z digest=sha256:20c688e0d5b006b0f0e48c51f415a4a26b19ce1e1916448e9ab65115f2fabca9

Observation 58a90809-e59c-4a97-b6b7-282250a13f89 · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:33:46.686607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T05:33:45.700040Z digest=sha256:d12ecb1279737708e2243ce7eef11a3c56656087618741988bcea7f190b7038f

Observation af77e99f-82b1-4031-b844-c52915452f42 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Does Self-Attention Need Separate Weights in Transformers? Generating Long Sequences with Sparse Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.712236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.712236Z digest=sha256:4a1457aaa8d247b9fb182da608283dd4e9beda04257f4a0f0742128da2f4ba3a

Observation 90c48c92-72e1-4185-b517-15d128710f3d · outbound

This paper cites Rethinking Attention with Performers.

Does Self-Attention Need Separate Weights in Transformers? Rethinking Attention with Performers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.723812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.723812Z digest=sha256:6ddb18ad1a0e0f127bee96f50f05b4f77d8f9d6deebe2097d51924203ab0fe23

Observation bb55ee82-d409-403a-9c24-1f05c717d4ad · outbound

This paper cites Symmetric Dot-Product Attention for Efficient Training of BERT Language Models.

Does Self-Attention Need Separate Weights in Transformers? Symmetric Dot-Product Attention for Efficient Training of BERT Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.734292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.734292Z digest=sha256:c94a5e85ac106b89e7b5efe272cd12915a79f2a32bf09d197568f1eb17d2c718

Observation 1ac29b27-c6f1-460e-8866-be2919825986 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Does Self-Attention Need Separate Weights in Transformers? BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.742845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.742845Z digest=sha256:1373f592b53a8beb3268e17e0cab7dc40f14374b7fa3a3edb596cc17a6eb891d

Observation 09015caf-434e-4037-81b5-d987ccdb2446 · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.753211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.753211Z digest=sha256:9fae0762892e9a1ce2ece299e127f29bd372a755539600c72aa93e70bc6e5f23

Observation 6fa064db-ba03-4ac2-9e47-06564cffb6a1 · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:33:46.655951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T05:33:45.760444Z digest=sha256:1c8b71d1c16a89962f56db0590fabe3703d2369a82459565492805a9a65f80d1

Observation 71e4f453-e78f-4868-b15d-563e6deecd5a · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.767177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.767177Z digest=sha256:f51d74493ca0f215ec761a295f9b088349e8201bd5560d6df0164e8cda3998e1

Observation bd4687e0-6f8f-4d16-a291-711d7de274d2 · outbound

This paper cites Simplifying Transformer Blocks.

Does Self-Attention Need Separate Weights in Transformers? Simplifying Transformer Blocks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.776668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.776668Z digest=sha256:cffcc4248634fb686fefa2317d143381c99a7dfd09103ff98067b93b9f855f6b

Observation 405c6c13-b43e-45ae-9245-b8c2ae6aee1d · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Does Self-Attention Need Separate Weights in Transformers? TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.789497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.789497Z digest=sha256:adf2724e84007c4e12a6e905f47a588c1fdc7be7a3a6b355a29eefea62946a3a

Observation 80bf6800-0e33-463e-bb53-55f598b96408 · outbound

This paper cites Exploring the Limits of Language Modeling.

Does Self-Attention Need Separate Weights in Transformers? Exploring the Limits of Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.796546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.796546Z digest=sha256:ed60b6d7daf660d024f069076fdbda0ebd2c067ab0af5e88330cc2c6ad818662

Observation 36a18289-d823-4b00-ab41-487e2b77c8b8 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Does Self-Attention Need Separate Weights in Transformers? Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.803817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.803817Z digest=sha256:7693aafbfee8cffc1457fa4a461ab549f008cf77cb831470bd4785bf19aeaa57

Observation 87629ccc-7007-4a1e-a68f-091a23f92a30 · outbound

This paper cites Reformer: The Efficient Transformer.

Does Self-Attention Need Separate Weights in Transformers? Reformer: The Efficient Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.810532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.810532Z digest=sha256:3b8bef36d3e388e3d2b1c7b39c082ea20724f686115ec875cda755ae665e90d7

Observation edfa42a1-5a15-41c0-9301-aea798b67cec · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:33:46.625672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T05:33:45.816579Z digest=sha256:29df6e9e6fff39af220ec8c356166e4b1712558563543c654a78850c15dbf153

Observation 7f6d6c27-8da7-4472-8947-81820ba9f4eb · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.831161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.831161Z digest=sha256:afaf214369964811bad05aaaeaf9ebe882f5b558c330b274b157e4f7438b90a1

Observation 73d58b2f-e0b9-46e9-a235-29f3efac9db1 · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:33:46.595494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T05:33:45.838607Z digest=sha256:8b65501da767f887e0490452e7a395d524f7d697eb84fe11d7fddf7a871cc6b0

Observation daaa3921-6c54-4992-bb9a-05e92964b19c · outbound

This paper cites Effective Approaches to Attention-based Neural Machine Translation.

Does Self-Attention Need Separate Weights in Transformers? Effective Approaches to Attention-based Neural Machine Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.847932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.847932Z digest=sha256:5ebbb2843630ef8f8c200ce1746eca489a42bbe0075273501699d489964adcdc

Observation ca0976db-88dd-4904-85d6-d972ae3ed60e · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.857160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.857160Z digest=sha256:9ecd3ff4185042c0efeb7094c2e68394e5ff37f7b5fa6a086ea390f2d98548f3

Observation 4b7966ec-184d-4ee2-aca1-aad55d9c8caf · outbound

This paper cites CoTexT: Multi-task Learning with Code-Text Transformer.

Does Self-Attention Need Separate Weights in Transformers? CoTexT: Multi-task Learning with Code-Text Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.864017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.864017Z digest=sha256:92f0c9a51f3c63ca6a7e0905c8f277f225b1a8232c7a79f54b1c4e07c1a49492

Observation 1f3b66f2-41c8-4b99-aa07-49c6aed6e2b2 · outbound

This paper cites Know What You Don't Know: Unanswerable Questions for SQuAD.

Does Self-Attention Need Separate Weights in Transformers? Know What You Don't Know: Unanswerable Questions for SQuAD

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.872936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.872936Z digest=sha256:6492af9225fbefb60ad85a13845a367d6058547d6dd444d28b1eb5adf68f300f

Observation eb2224f8-3102-4098-8730-2f0320c20c51 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Does Self-Attention Need Separate Weights in Transformers? SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.885909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.885909Z digest=sha256:94d1f1bd14bf21f3f99eefdd3de3c6a54740653d7238e79cb9feb1c3365fa10d

Observation bb3339a5-85a8-4761-8d5d-f13caf198849 · outbound

This paper cites Self-Attention with Relative Position Representations.

Does Self-Attention Need Separate Weights in Transformers? Self-Attention with Relative Position Representations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.892931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.892931Z digest=sha256:686017dddd48ecb7de49245083bf749ff0a8835ce1aeba9c03b9c0f828a238e6

Observation 0f3c06c1-c014-426d-ba6a-b05d5b20a182 · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.901138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.901138Z digest=sha256:c4e30e9ae702bb869e7646a1d56afedaee3fc8d3708f434c577d0595b1896e0e

Observation 56769c6a-04f2-4c63-8562-3a9ae09f4bed · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Does Self-Attention Need Separate Weights in Transformers? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.918943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.918943Z digest=sha256:3b968ffe8c8390e6260c38f23ccf906f9053118a21b0de475b138b0ec1875a5c

Observation 88ea327b-3e05-4ee5-81bd-fad767df586e · outbound

This paper cites RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder.

Does Self-Attention Need Separate Weights in Transformers? RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.925426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.925426Z digest=sha256:e3cd1e17bfccaec559e8f2dbac16ee525aaf1a91f0c562c52a807bed3d22339d

Observation d997a8c6-b468-4150-b079-e92eb652a042 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Does Self-Attention Need Separate Weights in Transformers? Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.932913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.932913Z digest=sha256:c42cb46c754229a86a5839c00158a2be74ad6c9d8ec2110a9f12ae8004db2569

Observation 1f4616fb-5c50-4ac9-b389-8fb4d3c925cc · outbound

This paper cites an unresolved cited work.

Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.940143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.940143Z digest=sha256:5fa8b464adbb3a58cb62ca5139a423d7326c2be720df43603213389bbac93ac6

Observation 2a3fb5e1-7586-498e-938d-2f761b859bdf · outbound

This paper cites A Survey on Efficient Training of Transformers.

Does Self-Attention Need Separate Weights in Transformers? A Survey on Efficient Training of Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.947692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.947692Z digest=sha256:8277c9318a266e17ec97053f05622f09e9fe241d344663f2330f9ed9cca8164b

Observation 0bec8e7d-0dc7-447b-b1b8-aeee4177e8ac · outbound

This paper cites online" 'onlinestring :=.

Does Self-Attention Need Separate Weights in Transformers? online" 'onlinestring :=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.969797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.969797Z digest=sha256:c2ceea1acd91f089aa68452def99da7021c6a27e9e07a536237936e238384b1e

Observation bb906104-cf6a-4864-aff5-ba7a2216aed8 · outbound

This paper cites write newline.

Does Self-Attention Need Separate Weights in Transformers? write newline

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:45.978048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:45.978048Z digest=sha256:1855a4d2c728ce4bc25c3ed4138f251877f24ad6533eb0648a180280927c9541

Pith citing papers

No inbound Pith citation observations are available.