Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:53:52.304925Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2412.18288.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:53:52.304925Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:05:02.502481Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T21:56:15.441801Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2f13138d-311a-40f9-8657-5c4d8e79a5a8 · outbound
Towards understanding how attention mechanism works in deep learning Proof: Denote the matrix of exp {fθ(xi, xj)} by W
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 138ac479-0c9f-4912-a910-d84cdb8ba9ba · outbound
Towards understanding how attention mechanism works in deep learning We conducted all experiments on a desktop computer with NVIDIA 2080Ti and 3.8 GHz AMD Ryzen 7 5800X 8-Core Process and 16 GB of memory and a computer with NVIDIA 3090Ti
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c84a1717-e171-4b54-a7a6-394f2527b763 · outbound
Towards understanding how attention mechanism works in deep learning A mathematical perspective on Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aef1b6c-3581-4434-bbbe-2b2ca271c7e7 · outbound
Towards understanding how attention mechanism works in deep learning UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd099c74-af2d-4578-beda-60b558401f39 · outbound
Towards understanding how attention mechanism works in deep learning Graph Attention Networks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e99909-dadb-4cf5-9784-4e12ffffaf4a · outbound
Towards understanding how attention mechanism works in deep learning URL http://dx.doi.org/10.18653/v1/2023.findings-acl
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4d7c30-78d5-4f86-ac62-e1cd680c8a91 · outbound
Towards understanding how attention mechanism works in deep learning Experiments and results T oy dataset We evaluated the performance of metric-attention by comparing it with self-attention and L2 self-attention using the Moon dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4508e9b1-f2a9-4128-b533-6fbaf5af89bd · outbound
Towards understanding how attention mechanism works in deep learning We adapted the implementation for the Multi30k dataset from https://github.com/hyunwoongko/transformer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bad485c5-fcff-4998-ae70-c902cac0c231 · outbound
Towards understanding how attention mechanism works in deep learning Manifold Fitting under Unbounded Noise
Reference 482
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 611a4699-9001-492a-b267-f481b1831bd0 · outbound
Towards understanding how attention mechanism works in deep learning Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space
Reference 1985
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbd2e32-7146-4b56-a5b2-1a9bf150283c · outbound
Towards understanding how attention mechanism works in deep learning doi: https://doi
Reference 1987
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c56f8573-1ad9-4b88-a1ef-60f37596e575 · outbound
Towards understanding how attention mechanism works in deep learning Score-Based Generative Modeling through Stochastic Differential Equations
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82944c5-dcb1-44ad-ae9b-4fc8c7682e65 · outbound
Towards understanding how attention mechanism works in deep learning A Mathematical Theory of Attention
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87404cd-b3f6-4a3a-bfdf-a0ad21ef621e · outbound
Towards understanding how attention mechanism works in deep learning Lan- guage models are few-shot learners
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a1cc6f57-d9fd-4fef-9c04-c32af807122d · outbound
Towards understanding how attention mechanism works in deep learning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b06985d4-25c0-4357-8667-253b52eac369 · outbound
Towards understanding how attention mechanism works in deep learning Scale-invariant heat kernel signatures for non- rigid shape recognition
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8d2806b5-461b-4c3a-8637-bdc8401350fd · outbound
Towards understanding how attention mechanism works in deep learning Multi30K: Multilingual English-German Image Descriptions
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefb950e-77b6-4e3f-ad29-dbdfc3a6a3f0 · outbound
Towards understanding how attention mechanism works in deep learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47f16bc-e2a8-4ea4-bb8e-8ec16c936997 · outbound
Towards understanding how attention mechanism works in deep learning Speech-transformer: A no-recurrence sequence- to-sequence model for speech recognition
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 114cfe83-b128-435b-9dc4-0ad77db2f974 · outbound
Towards understanding how attention mechanism works in deep learning Shape retrieval contest 2007: Wa- tertight models track
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe4b9dc8-3e46-43ed-9d3a-6f8ce3ca128e · inbound
Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs Towards understanding how attention mechanism works in deep learning
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f2568a-cf97-47ad-8fe9-c5fa3b1a24af · inbound
Attention's forward pass and Frank-Wolfe Towards understanding how attention mechanism works in deep learning
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5745ca25-4c14-4d87-9b02-61d40e725aa0 · inbound
Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation Towards understanding how attention mechanism works in deep learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf9ec9b5-68dd-413d-ac48-f943d92837ce · inbound
Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation Towards understanding how attention mechanism works in deep learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1740aba-fbd5-4a2a-bea8-58baccbf15cd · inbound
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference Towards understanding how attention mechanism works in deep learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4368ff3f-46d2-4554-b36c-331c141bbd05 · inbound
From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers Towards understanding how attention mechanism works in deep learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.