Pith. sign in

Paper Citation Record · LEDGER

Efficient Transformers: A Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2009.06732.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2009.06732 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:48.847581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:26.546053Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e99b873a-8ee7-47ea-bd68-6329a042f71d · inbound

Deformable DETR: Deformable Transformers for End-to-End Object Detection cites this paper.

Deformable DETR: Deformable Transformers for End-to-End Object Detection Efficient Transformers: A Survey

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:47:17.015597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T09:47:16.915936Z digest=sha256:a36c2cf693fd4fe0d0069d5e374e99da0ef2e2cbd4c8d70d1a816230886714c8

Observation 12edea71-84f4-4329-82d4-5de80f6a45f3 · inbound

Perceiver IO: A General Architecture for Structured Inputs & Outputs cites this paper.

Perceiver IO: A General Architecture for Structured Inputs & Outputs Efficient Transformers: A Survey

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:47:13.942394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T19:47:13.831505Z digest=sha256:888a9291d2f77499df4b5127a6187345f1833d2649aacc0ce50f4fd27c0691f2

Observation 7b2fb599-be62-4cc2-b64b-03d86bd32cef · inbound

Ligandformer: A Graph Neural Network for Predicting Compound Property with Robust Interpretation cites this paper.

Ligandformer: A Graph Neural Network for Predicting Compound Property with Robust Interpretation Efficient Transformers: A Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:24:27.496636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T12:22:23.989360Z digest=sha256:ecea2938de599e039791df6f2adca884443e6d106d26b7051ebf8668e1473aa5

Observation 5a3476db-bd8d-4fc1-9e49-5ef0d7a2f84f · inbound

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness cites this paper.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Efficient Transformers: A Survey

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:22:09.018695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T16:22:08.801066Z digest=sha256:3279de729e3cd225773ca6c9db39eabe81801db6ff7c3d286b44ae8ad8c51ded

Observation 4b4a1bac-4f88-491d-9e00-a224064295ab · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficient Transformers: A Survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.482223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:7ef571d9bd34158b449a05e5b1fa375ca4ed536f0bdb7701148a33e678cb2605

Observation 983eddb3-8be0-4bae-84c3-5ef51c7c26e1 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Efficient Transformers: A Survey

Reference 215

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:2b8d03ac2b96a818892a06f54e03ddf30b732bfdc808c54601d8218019095e7f

Observation b2c62249-e745-459e-b996-bb62726b4fed · inbound

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models cites this paper.

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models Efficient Transformers: A Survey

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:02.299100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:18:02.254491Z digest=sha256:1c7f6599e5064861abb23accd5f688f8e078b4f37dd47ecced675d1d9517d683

Observation 9800e6cd-56dc-449e-983a-e102a4bd1f38 · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Efficient Transformers: A Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.448737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:d4c890888067693a169db4f82966453d3fa4e1c287bd254ab32aab1a11e2538d

Observation 5de5ed35-0373-454a-a6b3-dcf414336252 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Efficient Transformers: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.285388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:e2051c4d887b2aa781e1b475d30720c630cb69d7b5adba276506dce7651c6ad1

Observation 4e49812c-ecd8-4cb3-817d-fde0fbffd9fd · inbound

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization cites this paper.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Efficient Transformers: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.847581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.847581Z digest=sha256:64d4ea37f92a540dc4488247bb70e4191f18da6a4c5296131b53cb5a337f4450

Observation 63bf97f3-c7f4-429d-99be-880d433f5eb0 · inbound

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer cites this paper.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient Transformers: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.842653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.842653Z digest=sha256:4800a42c5b81c923c59da0f59efb8db9f1a83cf9b9ccf2c9138f5da4dfa6f341

Observation 10f83ea2-3167-47af-be9d-9e55322ee048 · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Efficient Transformers: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:31.651037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:31.651037Z digest=sha256:34625f768986920787f7f9e6485af1705812a34c503ad1f95fc84fd58dda28ba

Observation 7fffd0f1-c1c2-4559-aff7-3ba47c27c151 · inbound

Crisp Attention: Regularizing Transformers via Structured Sparsity cites this paper.

Crisp Attention: Regularizing Transformers via Structured Sparsity Efficient Transformers: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:05.668093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:03:05.668093Z digest=sha256:b6f12c24b0c8f1124fa1f73d393ea88658b587be36861e46eeddec14d08cb679

Observation a5368d7b-d973-4ea7-865b-76c75cc6c8f4 · inbound

Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling cites this paper.

Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling Efficient Transformers: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:30:01.075667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:30:01.075667Z digest=sha256:d831383c2298c60536cb98d76396a91490fd8205750f1ec56bc4bacb29b44ae7

Observation b87d149a-da18-4137-9bc9-9bb63f236dd8 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Efficient Transformers: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.024663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.024663Z digest=sha256:96d3015972ba15235a1caa81330c13dbafcaa18115b873b3c1d79a1f7f862c6e

Observation 65c88820-912f-4ddd-9b9e-4559d0d1bf43 · inbound

ARC-Encoder: learning compressed text representations for large language models cites this paper.

ARC-Encoder: learning compressed text representations for large language models Efficient Transformers: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:28:52.455570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:28:52.455570Z digest=sha256:4b5e9af7601c36d5a7acec218535f6c2ad1edf676bc4003cfd0e470b0f45643c

Observation b693897c-c815-4367-a8f5-c0746837edd4 · inbound

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch cites this paper.

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch Efficient Transformers: A Survey

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:46:23.663710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:44:34.166709Z digest=sha256:e38dc24ee90d7e8586f4f1078a6d4fab47119f2832bd5469fb42cced534bc7de

Observation 35cf5da5-b212-4358-91a7-513006f24487 · inbound

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling cites this paper.

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling Efficient Transformers: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:39:59.305385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T11:35:43.088803Z digest=sha256:66878666c06d05a08bfeeebeaf4074def66bbbf97ada787c0814f1c14737017b

Observation b2158cce-f87d-465e-86a6-3cd321ac5793 · inbound

An explicit operator explains end-to-end computation in the modern neural networks used for sequence and language modeling cites this paper.

An explicit operator explains end-to-end computation in the modern neural networks used for sequence and language modeling Efficient Transformers: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:14.897597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:42:10.197470Z digest=sha256:bc7c9f320c93b5655331e71f75323d36eff657cbdd985e144194836f6eff4675

Observation 0b00fe47-86b9-4927-9a66-7b1ac81dd775 · inbound

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents cites this paper.

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents Efficient Transformers: A Survey

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:49:53.685071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T08:45:56.550821Z digest=sha256:cb1b2905de5deeec72de338ceca1cab0cd11060ce0c4e83bcef6936a4ec0f828

Observation adaaeadb-cf62-472d-858f-a58caed80e87 · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics Efficient Transformers: A Survey

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.547593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:0b15685a7657f7d84c4cd86eea3eb442c29c47c779b09edec83970eceffd402c

Observation 73cedafe-d713-46d4-875d-b82042af100b · inbound

TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings cites this paper.

TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings Efficient Transformers: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:52:55.820093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:52:55.820093Z digest=sha256:434fbb59034e135f146ad5f9c1cd395d6d7433998f48a33ccc07a773577f6be1