Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:33:45.978048Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2412.00359.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:33:45.978048Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dbcbe6c8-9a67-4f24-9517-9b468beac773 · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 556cc437-843c-49cf-be2c-e31782463e45 · outbound
Does Self-Attention Need Separate Weights in Transformers? Neural Machine Translation by Jointly Learning to Align and Translate
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5abf837-af36-4ddf-97f8-af1e1730064a · outbound
Does Self-Attention Need Separate Weights in Transformers? Longformer: The Long-Document Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a90809-e59c-4a97-b6b7-282250a13f89 · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af77e99f-82b1-4031-b844-c52915452f42 · outbound
Does Self-Attention Need Separate Weights in Transformers? Generating Long Sequences with Sparse Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c48c92-72e1-4185-b517-15d128710f3d · outbound
Does Self-Attention Need Separate Weights in Transformers? Rethinking Attention with Performers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb55ee82-d409-403a-9c24-1f05c717d4ad · outbound
Does Self-Attention Need Separate Weights in Transformers? Symmetric Dot-Product Attention for Efficient Training of BERT Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac29b27-c6f1-460e-8866-be2919825986 · outbound
Does Self-Attention Need Separate Weights in Transformers? BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09015caf-434e-4037-81b5-d987ccdb2446 · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa064db-ba03-4ac2-9e47-06564cffb6a1 · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71e4f453-e78f-4868-b15d-563e6deecd5a · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4687e0-6f8f-4d16-a291-711d7de274d2 · outbound
Does Self-Attention Need Separate Weights in Transformers? Simplifying Transformer Blocks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405c6c13-b43e-45ae-9245-b8c2ae6aee1d · outbound
Does Self-Attention Need Separate Weights in Transformers? TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80bf6800-0e33-463e-bb53-55f598b96408 · outbound
Does Self-Attention Need Separate Weights in Transformers? Exploring the Limits of Language Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a18289-d823-4b00-ab41-487e2b77c8b8 · outbound
Does Self-Attention Need Separate Weights in Transformers? Adam: A Method for Stochastic Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87629ccc-7007-4a1e-a68f-091a23f92a30 · outbound
Does Self-Attention Need Separate Weights in Transformers? Reformer: The Efficient Transformer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edfa42a1-5a15-41c0-9301-aea798b67cec · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f6d6c27-8da7-4472-8947-81820ba9f4eb · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d58b2f-e0b9-46e9-a235-29f3efac9db1 · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation daaa3921-6c54-4992-bb9a-05e92964b19c · outbound
Does Self-Attention Need Separate Weights in Transformers? Effective Approaches to Attention-based Neural Machine Translation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca0976db-88dd-4904-85d6-d972ae3ed60e · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b7966ec-184d-4ee2-aca1-aad55d9c8caf · outbound
Does Self-Attention Need Separate Weights in Transformers? CoTexT: Multi-task Learning with Code-Text Transformer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3b66f2-41c8-4b99-aa07-49c6aed6e2b2 · outbound
Does Self-Attention Need Separate Weights in Transformers? Know What You Don't Know: Unanswerable Questions for SQuAD
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2224f8-3102-4098-8730-2f0320c20c51 · outbound
Does Self-Attention Need Separate Weights in Transformers? SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb3339a5-85a8-4761-8d5d-f13caf198849 · outbound
Does Self-Attention Need Separate Weights in Transformers? Self-Attention with Relative Position Representations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3c06c1-c014-426d-ba6a-b05d5b20a182 · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56769c6a-04f2-4c63-8562-3a9ae09f4bed · outbound
Does Self-Attention Need Separate Weights in Transformers? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ea327b-3e05-4ee5-81bd-fad767df586e · outbound
Does Self-Attention Need Separate Weights in Transformers? RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d997a8c6-b468-4150-b079-e92eb652a042 · outbound
Does Self-Attention Need Separate Weights in Transformers? Deformable DETR: Deformable Transformers for End-to-End Object Detection
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4616fb-5c50-4ac9-b389-8fb4d3c925cc · outbound
Does Self-Attention Need Separate Weights in Transformers? Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3fb5e1-7586-498e-938d-2f761b859bdf · outbound
Does Self-Attention Need Separate Weights in Transformers? A Survey on Efficient Training of Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bec8e7d-0dc7-447b-b1b8-aeee4177e8ac · outbound
Does Self-Attention Need Separate Weights in Transformers? online" 'onlinestring :=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb906104-cf6a-4864-aff5-ba7a2216aed8 · outbound
Does Self-Attention Need Separate Weights in Transformers? write newline
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.