Pith. sign in

Paper Citation Record · LEDGER

A Survey on Transformer Compression

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.05964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05964 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:41:21.076278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:17.329300Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e5340ae3-5225-4259-a177-bacab80a4daf · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models A Survey on Transformer Compression

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.600583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:d66a58e457be301b61134b2f368bf8418987c740b558e7baeea5906d71f5a5b2

Observation d9a3526e-77da-4b6e-9fbe-157d4cb75c35 · inbound

A Comparative Study of Pruning Methods in Transformer-based Time Series Forecasting cites this paper.

A Comparative Study of Pruning Methods in Transformer-based Time Series Forecasting A Survey on Transformer Compression

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:41:21.076278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:41:21.076278Z digest=sha256:f1d4fba3ed6e50cac54d3bcba122ec039005d7cd3ac547c0126cdd441e892f73

Observation bde524e4-757f-491a-8c72-fb587c7ce792 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management A Survey on Transformer Compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.589713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.589713Z digest=sha256:cfebed8dac5c897bb5d96eb73d55617e5eaf342cc2c4dc6be05bffffe0e0e636

Observation a8167720-6f32-4f08-a8fe-57a3f9f7af08 · inbound

Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models cites this paper.

Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models A Survey on Transformer Compression

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:34.067027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:27:59.317818Z digest=sha256:6ccda5391daf851ec944e18751aee4fb638464100b797550ac918fc6c13bc8d9

Observation dcd730a1-9d8a-4eb8-909d-20fbf5061899 · inbound

Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones cites this paper.

Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones A Survey on Transformer Compression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:23.468653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:23.468653Z digest=sha256:8568f5c3574e17ea379fc333fb424bfc99821a44072f9bbe8925f3d73a161e69

Observation f9663580-685a-44e5-9093-ea1f5b4e7451 · inbound

DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs cites this paper.

DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs A Survey on Transformer Compression

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:59:20.593821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:59:20.593821Z digest=sha256:2b677dbfd7a4927f50490e24c5f5da6d90b92d3f8484e7f7f24faafc53543d2c

Observation b99fc76a-9cde-4e3d-b033-a7d0505ff449 · inbound

LinDeps: A Fine-tuning Free Post-Pruning Method to Remove Layer-Wise Linear Dependencies with Guaranteed Performance Preservation cites this paper.

LinDeps: A Fine-tuning Free Post-Pruning Method to Remove Layer-Wise Linear Dependencies with Guaranteed Performance Preservation A Survey on Transformer Compression

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:05.892391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:05.892391Z digest=sha256:9dd3926ce8c283491dee3b8a889f22ac1ff293ee7bef2f430c26e49fd64d53ab

Observation 66aa30a7-72a9-4c8d-bd1d-bdd1aa3a9284 · inbound

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models cites this paper.

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models A Survey on Transformer Compression

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:44:26.532491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T23:44:01.953344Z digest=sha256:66e6e78b3eda0ccfde7ea778c0a08228e40cf5f7bd9865bba1bdfbffc585a319

Observation c5f89143-76b1-4882-abdf-869b49aab718 · inbound

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization cites this paper.

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization A Survey on Transformer Compression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T13:12:14.563587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:12:14.563587Z digest=sha256:8d9d150b20d3c8da049e908fef7ca7635af15fd209ca50d630dc17d04569cfc4

Observation e89911bd-6ab5-478a-8be7-90820dc3534d · inbound

Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction cites this paper.

Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction A Survey on Transformer Compression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T10:12:43.502045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:12:43.502045Z digest=sha256:d1b741280164c5e18a415721686263aaebe30273680c31f563ac2a35c35b6414

Observation 8240d2a2-a042-43a1-9d48-a0ae110b3cc1 · inbound

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs cites this paper.

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs A Survey on Transformer Compression

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.951036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T06:33:14.766545Z digest=sha256:1fe268fbc68cfc34ce77b08a48b00923e9e6875f13fc157c9fe313a3ada4f715

Observation d403bebb-903f-4fb2-b790-235b821540f0 · inbound

Depth Adaptive Efficient Visual Autoregressive Modeling cites this paper.

Depth Adaptive Efficient Visual Autoregressive Modeling A Survey on Transformer Compression

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.622851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T06:22:55.035749Z digest=sha256:52914854a1489164f3316d7c2c9af190f1283602f3fec551999e47bfe158ad45

Observation 1e29b0d3-ed6f-48e1-91d7-e77941baf7c3 · inbound

Do Transformers Need Three Projections? Systematic Study of QKV Variants cites this paper.

Do Transformers Need Three Projections? Systematic Study of QKV Variants A Survey on Transformer Compression

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:17.330967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T15:14:49.475561Z digest=sha256:6054959f91242ea69645c346cef59d850b4291592fa94fd88d43d7254d2e4de0

Observation 3f32ec1b-36c5-44d4-b315-a9b0c5d86c85 · inbound

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers cites this paper.

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers A Survey on Transformer Compression

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T02:57:29.274601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:57:29.274601Z digest=sha256:5ab4b504567e0e5c871dce90477885631224afb3ee9f313ac336946a239c4b35