Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:36:28.735305Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2505.18413.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:36:28.735305Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2c06224-b566-46f9-b865-6b7bff397c7b · outbound
LatentLLM: Attention-Aware Joint Tensor Compression GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1deea97-6f15-4f39-bd23-c5e60d46039f · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db0197c-8592-47b0-aec6-67b04e4343ce · outbound
LatentLLM: Attention-Aware Joint Tensor Compression SparseLLM: Towards global pruning of pre-trained language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c685ec87-3a1c-4d35-81be-3caa902f11fe · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de377a09-f03b-457f-b1bd-0eb0dffa5fe3 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Palu: Compressing KV-Cache with Low-Rank Projection
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 240a159d-da0c-4316-a833-afe69cfeaab9 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f4e7c6-8291-44ca-814d-f623ae7e283f · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 17241804-7cb0-4bec-a9e8-0a2140aa1046 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Exploiting linear structure within convolutional networks for efficient evaluation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bac52161-ec82-4edd-92f7-0ab3091b3fd8 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression The case for 4-bit pre- cision: k-bit inference scaling laws
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0ea253db-e85d-4860-8d87-befe9b60dc66 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression SparseGPT: Massive lan- guage models can be accurately pruned in one-shot
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b19b3af-67d6-45e6-a5cb-488ffefaae87 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4e12b9-8b99-4d64-ad86-811e24049369 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Optimal brain surgeon and general network pruning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8504ecd0-91f8-41a5-be47-075d2c859709 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14c94767-f9f5-4430-826a-5038496469c3 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8671d7-b1bf-4560-af95-bd7dceed05d2 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Mixtral of Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f00abf-534c-48e9-ad7d-16d00534f531 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression GPT-4 passes the bar exam
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8145a8c0-1983-4bed-8c60-c98ed18a1a49 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression BERT: Pre-training of deep bidirectional trans- formers for language understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7bd49f21-f817-4a6b-bc8a-ebab02acf37c · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Optimal brain damage
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1071d354-ec1c-47bb-94bb-f113775a96b0 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression A well-conditioned esti- mator for large-dimensional covariance matrices
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 583760f6-b70a-42c5-9c3a-8afa8482b7b6 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression LoSparse: Structured com- pression of large language models based on low-rank and sparse approximation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6740093-4468-4bfb-98e9-ea5c20b7031a · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ebe3966-64f4-492c-a488-b3e3cf0c3119 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ad3544-b528-4917-b96a-df0dcc96819b · outbound
LatentLLM: Attention-Aware Joint Tensor Compression AWQ: Activation-aware weight quantization for on-device LLM compression and accelera- tion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8284d0aa-2043-4fd3-a339-1ebdb6dd188f · outbound
LatentLLM: Attention-Aware Joint Tensor Compression DeepSeek-V3 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e634fac1-8581-4ad3-b278-992e3eb03c67 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Visual instruction tuning, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d47d09b3-8be8-4d51-87c8-b0a01a197364 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 932a920d-14f5-4e6d-9c14-3e22c81754f0 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression The penn treebank: Annotating pred- icate argument structure
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf4e8bf6-d8af-4ab3-bd6f-b2f5d8778f46 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Pointer Sentinel Mixture Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47530378-0076-485c-a6a3-5bafe57adec1 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression PyTorch: An imperative style, high-performance deep learning li- brary
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e5a97581-de30-45d0-8164-2ca0c1fb69ae · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Improving language understanding by gener- ative pre-training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ded2fb4a-2627-4b35-b587-87a07647adc0 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 308176fb-f749-456e-b4df-b5e993404494 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Compressing large language models using low rank and low precision decomposition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 81cb0c3b-46dc-41a3-88f7-f9d8c5d4fcb3 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Low-rank matrix factorization for deep neural network training with high- dimensional output targets
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 508cc0d9-098f-4f60-b8ab-01ccc9a87c36 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20dda85b-cc72-4e9e-9352-e82849ede1a5 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Low-rank lottery tick- ets: finding efficient low-rank neural networks via matrix differential equations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2ec3fa6e-c35f-458c-951c-7b43c5e80af2 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Green AI
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e7a8c2e-c841-4268-b95a-cabab14488f8 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression RoFormer: Enhanced transformer with rotary position embedding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18f6aca1-b490-41b1-aa38-f50ef6ff4c85 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression A Simple and Effective Pruning Approach for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c43f400-fd57-4aa1-8161-ce727b13c736 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76d86c79-9ba3-471e-ae7f-c9e467fe44df · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Q-VLM: Post-training Quantization for Large Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603ecd6f-dbb6-4f96-bcc4-d59ffd567631 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9bb04e-07e6-4706-9bff-98ca4f8da6a1 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression Emergent Abilities of Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34cbf179-46fa-4b85-84a3-6ce0b858f452 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb1f7871-a812-4f81-ac68-d52ca000203e · outbound
LatentLLM: Attention-Aware Joint Tensor Compression A survey on model com- pression and acceleration for pretrained language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 657eba76-ddd8-4c71-9aa4-d7ed7b2be546 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9657a837-80cd-4eb0-91fc-86a062381182 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression ZeroQuant: Ef- ficient and affordable post-training quantization for large- scale transformers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58b88554-8138-461b-adcc-18dfbad1c2d7 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b93a57aa-0a94-4728-8bc8-8f5214d49e12 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e09b5b-3bc5-445f-a80e-64b3cbf75517 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression OPT: Open Pre-trained Transformer Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242077bb-9032-41fe-a6f4-4057b8020261 · outbound
LatentLLM: Attention-Aware Joint Tensor Compression C 1 2 O µ⊤C −1 2 (1 − µ⊤C +µ) 1 2 #
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 235a3222-a37a-4581-91b0-6192fddafa9c · outbound
LatentLLM: Attention-Aware Joint Tensor Compression (193) Plugging into the loss gives: L = X i ∥Wo,iWv,i(X − µ1⊤) − ˆWo,i ˆWv,i(X − µ1⊤)∥2 (194) = X i ∥ Wo,iWv,i| {z } Gi∈Rd×d C 1 2 0 − Bo Ao,iBv,i| {z } Hi∈Rro ×rv AvC 1 2 0 ∥2
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.