Pith. sign in

Paper Citation Record · LEDGER

Full Stack Optimization of Transformer Inference: a Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2302.14017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.14017 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.547913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:38:49.031648Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1dce0f8e-e1f0-4d76-84ba-eaf82b5e4aa3 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Full Stack Optimization of Transformer Inference: a Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.547913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.547913Z digest=sha256:8d583487d4e790bfa5cda2b137a572143e10f713f04207a99910622251d7c96d

Observation 295d31ee-b35e-4921-8889-44aef5a7985c · inbound

A distillation-teleportation protocol for fault-tolerant QRAM cites this paper.

A distillation-teleportation protocol for fault-tolerant QRAM Full Stack Optimization of Transformer Inference: a Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:55.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:55.019799Z digest=sha256:051c75b074dc476b9bcd067bdec46d1175c93c0b6920415fa951fece3267e031

Observation ff7dada2-ed6b-4b18-94ab-09072f391012 · inbound

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations cites this paper.

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations Full Stack Optimization of Transformer Inference: a Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:28.605678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:28.605678Z digest=sha256:3fc83455f3c026d57e5598983d378d321ffccd6dca1cac70e810a867c95d883f

Observation 250be190-98cb-4066-b14f-2ead6dde0bdb · inbound

COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives cites this paper.

COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives Full Stack Optimization of Transformer Inference: a Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:29:57.252635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:29:57.252635Z digest=sha256:5ed108707d2dfcb912209c0c6d34b9941698d36ed61a6cfaa39495fff6b5d90b

Observation 899656f8-c19e-41f3-aed5-d9ee1e929726 · inbound

vAttention: Verified Sparse Attention cites this paper.

vAttention: Verified Sparse Attention Full Stack Optimization of Transformer Inference: a Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:21:07.403398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:21:07.403398Z digest=sha256:a8dca00f74ac5fdc16038a7239ee819282944a19a41e00879eba0d79cb0c8660

Observation de6ce19a-7c98-411b-843c-8cf7c0e94d6d · inbound

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs cites this paper.

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs Full Stack Optimization of Transformer Inference: a Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.960085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:33:14.766545Z digest=sha256:ccc3144ce67887fd1fc89c75beb077cea08b891d07277494d8cf082b19046a97

Observation 8cb40e57-c913-4f11-b89c-d47ac3da4d64 · inbound

Learning to Remember, Learn, and Forget in Attention-Based Models cites this paper.

Learning to Remember, Learn, and Forget in Attention-Based Models Full Stack Optimization of Transformer Inference: a Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:06.492108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:06.492108Z digest=sha256:83128f3c5ab2ad62ec41c6e3bc97154f9656df4ff305af0132782b35604cc632

Observation 6c75d491-37a0-43e5-a01f-73441826f58a · inbound

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures cites this paper.

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures Full Stack Optimization of Transformer Inference: a Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.486747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:33:59.777818Z digest=sha256:ceeec0c71e8552e559933482c98006f038e2806ccdc82b902cf2a9bdbb71fb18

Observation af522d9d-bae7-4c65-94db-fd8fefdb935a · inbound

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation cites this paper.

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation Full Stack Optimization of Transformer Inference: a Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T23:32:59.467653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:32:59.467653Z digest=sha256:bab8b4c67d873dc676b5f304a8225fcae43b52344c38524bfc74cae42bdb6289

Observation 26e3db36-ea6b-4493-bfe9-30bd271a7079 · inbound

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models cites this paper.

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models Full Stack Optimization of Transformer Inference: a Survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:16:01.588643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:04:49.793158Z digest=sha256:faa5c597c4ee3d3c74a2f344c5c0457f260b0286324e4582b7661a2deb1b556b

Observation 431c1525-8904-44cc-9e97-8d137d5a0425 · inbound

CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration cites this paper.

CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration Full Stack Optimization of Transformer Inference: a Survey

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.117402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:56:53.425178Z digest=sha256:0d428d9a4c62b9cecfac9c1acb0b453b23be76a166052b83a7665d8682f44208

Observation 9e9887d7-95d9-4a33-a239-09e47bd96931 · inbound

Edge-Inference Governors Need Memory-Clock State cites this paper.

Edge-Inference Governors Need Memory-Clock State Full Stack Optimization of Transformer Inference: a Survey

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:38:49.033705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T02:44:49.539266Z digest=sha256:5d6962de66e6dfb7df6ab2b6317b40a049c4a98dccd9df3f9debe13cd7cf20e3

Observation 5fd2401b-6fbc-46ee-b85b-96ae1df2ebdc · inbound

Edge-Inference Governors Need Memory-Clock State cites this paper.

Edge-Inference Governors Need Memory-Clock State Full Stack Optimization of Transformer Inference: a Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T13:55:19.148570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:55:19.148570Z digest=sha256:5ad231131d5576ae042a29e5f4ead782e4d9da1b05b6410dc36b7e8ac5256aae