Pith. sign in

Paper Citation Record · LEDGER

Towards understanding how attention mechanism works in deep learning

As of 12 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 6 inbound Pith citation observations for arXiv:2412.18288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18288 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:53:52.304925Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:05:02.502481Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:56:15.441801Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f13138d-311a-40f9-8657-5c4d8e79a5a8 · outbound

This paper cites Proof: Denote the matrix of exp {fθ(xi, xj)} by W.

Towards understanding how attention mechanism works in deep learning Proof: Denote the matrix of exp {fθ(xi, xj)} by W

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.621405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.304925Z digest=sha256:8cc735dcdf89afef8c6078306344fd55e3c4750ae41dd831d5347ff8bd3a72c7

Observation 138ac479-0c9f-4912-a910-d84cdb8ba9ba · outbound

This paper cites We conducted all experiments on a desktop computer with NVIDIA 2080Ti and 3.8 GHz AMD Ryzen 7 5800X 8-Core Process and 16 GB of memory and a computer with NVIDIA 3090Ti.

Towards understanding how attention mechanism works in deep learning We conducted all experiments on a desktop computer with NVIDIA 2080Ti and 3.8 GHz AMD Ryzen 7 5800X 8-Core Process and 16 GB of memory and a computer with NVIDIA 3090Ti

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.651971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.295114Z digest=sha256:a160da1ef405c064d15b399f5ca4527c0b0ff69c45a94673c967973f3234e10e

Observation c84a1717-e171-4b54-a7a6-394f2527b763 · outbound

This paper cites A mathematical perspective on Transformers.

Towards understanding how attention mechanism works in deep learning A mathematical perspective on Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.247956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.247956Z digest=sha256:753fa3c3fa125cbd085d8e8bc0a95e7c60c50aa87ceb271e52a346c549f55d12

Observation 1aef1b6c-3581-4434-bbbe-2b2ca271c7e7 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Towards understanding how attention mechanism works in deep learning UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.257258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.257258Z digest=sha256:58f4ab1502fc8c544db978687a2e9637e82264ef6c3ed9f5df345e7e46ccfd91

Observation bd099c74-af2d-4578-beda-60b558401f39 · outbound

This paper cites Graph Attention Networks.

Towards understanding how attention mechanism works in deep learning Graph Attention Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.267529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.267529Z digest=sha256:760141acd752129379ee8dee5956352da2505a876f6d9351c2c5ec8ac49949c2

Observation d0e99909-dadb-4cf5-9784-4e12ffffaf4a · outbound

This paper cites URL http://dx.doi.org/10.18653/v1/2023.findings-acl.

Towards understanding how attention mechanism works in deep learning URL http://dx.doi.org/10.18653/v1/2023.findings-acl

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-11T04:53:52.281360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.281360Z digest=sha256:6d341bdfe79576384b39afdf0bb7f609f7220d416eca3f5ba984964fd3fd15f8

Observation 9a4d7c30-78d5-4f86-ac62-e1cd680c8a91 · outbound

This paper cites Experiments and results T oy dataset We evaluated the performance of metric-attention by comparing it with self-attention and L2 self-attention using the Moon dataset.

Towards understanding how attention mechanism works in deep learning Experiments and results T oy dataset We evaluated the performance of metric-attention by comparing it with self-attention and L2 self-attention using the Moon dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.668163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.290604Z digest=sha256:01df87e41895442bf6858b49df6c2e9eae916219edbb0bbe6c8da5db96c31371

Observation 4508e9b1-f2a9-4128-b533-6fbaf5af89bd · outbound

This paper cites We adapted the implementation for the Multi30k dataset from https://github.com/hyunwoongko/transformer.

Towards understanding how attention mechanism works in deep learning We adapted the implementation for the Multi30k dataset from https://github.com/hyunwoongko/transformer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.635416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.299600Z digest=sha256:3141fc5b9d0111df67c2800560b78c0f2249c8a3d9c3be4431b0b90424670e10

Observation bad485c5-fcff-4998-ae70-c902cac0c231 · outbound

This paper cites Manifold Fitting under Unbounded Noise.

Towards understanding how attention mechanism works in deep learning Manifold Fitting under Unbounded Noise

Reference 482

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:53:52.373924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.285731Z digest=sha256:2f5102fd531a5b5e2f25cf5a4cd73961f72c7a51f3a9b516436b1ac73876edd4

Observation 611a4699-9001-492a-b267-f481b1831bd0 · outbound

This paper cites Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space.

Towards understanding how attention mechanism works in deep learning Deep Residual Networks Learn the Geodesic Curve in the Wasserstein Space

Reference 1985

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.243199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.243199Z digest=sha256:34d7440837a00d0707f1e30682f1bb0bf10e295d4c727acdab91faf8fc7cbd92

Observation ddbd2e32-7146-4b56-a5b2-1a9bf150283c · outbound

This paper cites doi: https://doi.

Towards understanding how attention mechanism works in deep learning doi: https://doi

Reference 1987

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.277386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.277386Z digest=sha256:7c000a1037f33a4a5034e7ca9cb8b8ce35078f64d8db33ca8fb44b2f7a4f472b

Observation c56f8573-1ad9-4b88-a1ef-60f37596e575 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Towards understanding how attention mechanism works in deep learning Score-Based Generative Modeling through Stochastic Differential Equations

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.262557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.262557Z digest=sha256:9f5c902f31277c53b16c0ac9517c1aadb16a905ebd10a6daba30b6f50cceb0e4

Observation a82944c5-dcb1-44ad-ae9b-4fc8c7682e65 · outbound

This paper cites A Mathematical Theory of Attention.

Towards understanding how attention mechanism works in deep learning A Mathematical Theory of Attention

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.272809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.272809Z digest=sha256:19e4e25d5b24d31735eb579b3e3eab9aedee5280aca0d2e0f8c5efc869477b6f

Observation e87404cd-b3f6-4a3a-bfdf-a0ad21ef621e · outbound

This paper cites Lan- guage models are few-shot learners.

Towards understanding how attention mechanism works in deep learning Lan- guage models are few-shot learners

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.720224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.215981Z digest=sha256:7d454f3a066c5780d8f2d3ba3ae2bafe9c6d8942fe0068044a3220d761b9e261

Observation a1cc6f57-d9fd-4fef-9c04-c32af807122d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Towards understanding how attention mechanism works in deep learning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.221496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.221496Z digest=sha256:dea7cf8b699eee7a59234e9ba1d07bf8eba5231a2782f91a6c1c2616b17d81bb

Observation b06985d4-25c0-4357-8667-253b52eac369 · outbound

This paper cites Scale-invariant heat kernel signatures for non- rigid shape recognition.

Towards understanding how attention mechanism works in deep learning Scale-invariant heat kernel signatures for non- rigid shape recognition

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.733661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.210049Z digest=sha256:51e6a3b79d423eefc1f644b6472cd62fc998ffcffd665544815d8755605ab288

Observation 8d2806b5-461b-4c3a-8637-bdc8401350fd · outbound

This paper cites Multi30K: Multilingual English-German Image Descriptions.

Towards understanding how attention mechanism works in deep learning Multi30K: Multilingual English-German Image Descriptions

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.237619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.237619Z digest=sha256:7f1d287de371f91a4db1ab0c2379c4a9b6ac487511465bacedec0e75dd43a6b5

Observation cefb950e-77b6-4e3f-ad29-dbdfc3a6a3f0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Towards understanding how attention mechanism works in deep learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:52.232364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:52.232364Z digest=sha256:79b4b48bfbf514a5c019a9c89a10c23e60fad555c565f82f28634c73ef422775

Observation e47f16bc-e2a8-4ea4-bb8e-8ec16c936997 · outbound

This paper cites Speech-transformer: A no-recurrence sequence- to-sequence model for speech recognition.

Towards understanding how attention mechanism works in deep learning Speech-transformer: A no-recurrence sequence- to-sequence model for speech recognition

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.703501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.227828Z digest=sha256:7e994592658d5e52f0b7f9f864aa9cd8c7889c0b65654ef2da87bd88bec8910d

Observation 114cfe83-b128-435b-9dc4-0ad77db2f974 · outbound

This paper cites Shape retrieval contest 2007: Wa- tertight models track.

Towards understanding how attention mechanism works in deep learning Shape retrieval contest 2007: Wa- tertight models track

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:53:52.684496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:53:52.252781Z digest=sha256:55f8bf640c7e49e430af6f2fef6ab017c790dc3ad5248716fe7dfd443a52cfcd

Pith citing papers

Observation fe4b9dc8-3e46-43ed-9d3a-6f8ce3ca128e · inbound

Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs cites this paper.

Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs Towards understanding how attention mechanism works in deep learning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-05T20:59:36.897052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:59:36.897052Z digest=sha256:600f5021ed57c1333e030c340b0412e0a45c44aeb70bfcba913e9ec6d3f41a7e

Observation 02f2568a-cf97-47ad-8fe9-c5fa3b1a24af · inbound

Attention's forward pass and Frank-Wolfe cites this paper.

Attention's forward pass and Frank-Wolfe Towards understanding how attention mechanism works in deep learning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-05T21:05:02.502481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:05:02.502481Z digest=sha256:59b7937c3e3ce5e3dc35b43db7f9280c2b3d6e7142d0e88ae3900efc19f70439

Observation 5745ca25-4c14-4d87-9b02-61d40e725aa0 · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation Towards understanding how attention mechanism works in deep learning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:01:46.313882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T18:56:48.722344Z digest=sha256:0d4054d7170cd2d73106894f735cfb6b809d9dc1b13ad491ef46cabd87c1dd24

Observation bf9ec9b5-68dd-413d-ac48-f943d92837ce · inbound

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation cites this paper.

Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation Towards understanding how attention mechanism works in deep learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T10:25:17.199621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:25:17.199621Z digest=sha256:b11abd9834f0b088282d65d04f0029e21c53a995fe550e65b2a20e38f3c9e460

Observation d1740aba-fbd5-4a2a-bea8-58baccbf15cd · inbound

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference cites this paper.

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference Towards understanding how attention mechanism works in deep learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.443916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T16:07:22.196982Z digest=sha256:715dbae7f1a5c44649a07f30202da24128db2d8c0c0a0ffefc8bc38b7b751aea

Observation 4368ff3f-46d2-4554-b36c-331c141bbd05 · inbound

From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers cites this paper.

From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers Towards understanding how attention mechanism works in deep learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T10:03:27.747512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:03:27.747512Z digest=sha256:203782b00f06e985da54b40e8bcee83d38ebda93cea802f376dc6899858ca4f3