Pith. sign in

Paper Citation Record · LEDGER

LatentLLM: Attention-Aware Joint Tensor Compression

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2505.18413.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18413 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:36:28.735305Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2c06224-b566-46f9-b865-6b7bff397c7b · outbound

This paper cites GPT-4 Technical Report.

LatentLLM: Attention-Aware Joint Tensor Compression GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.282246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.282246Z digest=sha256:5895431e51385fd6cf5c2ddf78ab93d1e62f172333aeb788368de6f5383a49a9

Observation d1deea97-6f15-4f39-bd23-c5e60d46039f · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.329311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.329311Z digest=sha256:5a85260b7ef86eaba2901fc85ca431f87be7cc5411dbdc62b2a3fbbc242b203f

Observation 1db0197c-8592-47b0-aec6-67b04e4343ce · outbound

This paper cites SparseLLM: Towards global pruning of pre-trained language models.

LatentLLM: Attention-Aware Joint Tensor Compression SparseLLM: Towards global pruning of pre-trained language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:32.810350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:25.417910Z digest=sha256:205a47638b66b2a6ac8a66f96565e99aa0cbec666ac65b555bb72c413e93fa92

Observation c685ec87-3a1c-4d35-81be-3caa902f11fe · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

LatentLLM: Attention-Aware Joint Tensor Compression Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.480363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.480363Z digest=sha256:d3fd9e845292509330663cb0f2dcfb8c4dc493ff28b9bd75a537a324330140a6

Observation de377a09-f03b-457f-b1bd-0eb0dffa5fe3 · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

LatentLLM: Attention-Aware Joint Tensor Compression Palu: Compressing KV-Cache with Low-Rank Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.552150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.552150Z digest=sha256:58ac6f2dc0129d4090b1b8508851bdede849e55597415a57cd8695b9d86d9e07

Observation 240a159d-da0c-4316-a833-afe69cfeaab9 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

LatentLLM: Attention-Aware Joint Tensor Compression Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.615439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.615439Z digest=sha256:811e424e9f0b852b118cb943a228ec8a67591b0fabe8bfd3f8ae9aa28c5c4214

Observation 11f4e7c6-8291-44ca-814d-f623ae7e283f · outbound

This paper cites an unresolved cited work.

LatentLLM: Attention-Aware Joint Tensor Compression Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:36:32.681512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:25.680177Z digest=sha256:370730c3befd566dbe53ae361367d57d806b22188f9df896efc4dc5578d9f5a4

Observation 17241804-7cb0-4bec-a9e8-0a2140aa1046 · outbound

This paper cites Exploiting linear structure within convolutional networks for efficient evaluation.

LatentLLM: Attention-Aware Joint Tensor Compression Exploiting linear structure within convolutional networks for efficient evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:32.559494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:25.755262Z digest=sha256:4d3bd910fc2742cba969427b2ac5707b7ccb842233649f0b2882b9d44af5249f

Observation bac52161-ec82-4edd-92f7-0ab3091b3fd8 · outbound

This paper cites The case for 4-bit pre- cision: k-bit inference scaling laws.

LatentLLM: Attention-Aware Joint Tensor Compression The case for 4-bit pre- cision: k-bit inference scaling laws

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:32.368879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:25.838247Z digest=sha256:afc5c0082e836a378696c925009c95b7003f895925f28618caf8d3e43a9f80ab

Observation 0ea253db-e85d-4860-8d87-befe9b60dc66 · outbound

This paper cites SparseGPT: Massive lan- guage models can be accurately pruned in one-shot.

LatentLLM: Attention-Aware Joint Tensor Compression SparseGPT: Massive lan- guage models can be accurately pruned in one-shot

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:32.224670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:25.901574Z digest=sha256:1eacc2671772bd94a6eed54adbdbad2d603be9bb7a6c63c73e18a493b207b0f8

Observation 0b19b3af-67d6-45e6-a5cb-488ffefaae87 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

LatentLLM: Attention-Aware Joint Tensor Compression GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:25.988734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:25.988734Z digest=sha256:901ec11734edd4fa5d318ec6ebfd7cb9a2cd5688770319892895bc0842a45eec

Observation 7d4e12b9-8b99-4d64-ad86-811e24049369 · outbound

This paper cites Optimal brain surgeon and general network pruning.

LatentLLM: Attention-Aware Joint Tensor Compression Optimal brain surgeon and general network pruning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:32.103562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.071748Z digest=sha256:61eb83307bf375215e2fbd385a05ddbcfe4ecbed9109fbc8a881ba6456e68b1e

Observation 8504ecd0-91f8-41a5-be47-075d2c859709 · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

LatentLLM: Attention-Aware Joint Tensor Compression Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:26.140525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:26.140525Z digest=sha256:c604953858616fb97a9be87bbcbf97a6d3b76fcbe45dba68ab4f62014af5eba8

Observation 14c94767-f9f5-4430-826a-5038496469c3 · outbound

This paper cites PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation.

LatentLLM: Attention-Aware Joint Tensor Compression PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:26.200703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:26.200703Z digest=sha256:159f929944ead7e15b9334198fa6ac8e77372a50a590a27d3236b9d58b51cb35

Observation 0d8671d7-b1bf-4560-af95-bd7dceed05d2 · outbound

This paper cites Mixtral of Experts.

LatentLLM: Attention-Aware Joint Tensor Compression Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:26.273603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:26.273603Z digest=sha256:098139a04e1f40d9024ac474c63eddf8ce4d586bf47fbc3a11845abc506dde38

Observation 77f00abf-534c-48e9-ad7d-16d00534f531 · outbound

This paper cites GPT-4 passes the bar exam.

LatentLLM: Attention-Aware Joint Tensor Compression GPT-4 passes the bar exam

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.978651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.350686Z digest=sha256:a0042f3770fe7f0f6104918f372d85e384f714b61b9001378b75121e99c127a3

Observation 8145a8c0-1983-4bed-8c60-c98ed18a1a49 · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

LatentLLM: Attention-Aware Joint Tensor Compression BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.835035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.436078Z digest=sha256:69330fd472dbcf3f3e925cb6a47e5c872fa6913ce7521b14c914eedf509cbbf3

Observation 7bd49f21-f817-4a6b-bc8a-ebab02acf37c · outbound

This paper cites Optimal brain damage.

LatentLLM: Attention-Aware Joint Tensor Compression Optimal brain damage

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.720301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.521191Z digest=sha256:7bb87a3536f3ec8521017a4f891d1f87bc99af66c11c7adf0292b7e789d6fab6

Observation 1071d354-ec1c-47bb-94bb-f113775a96b0 · outbound

This paper cites A well-conditioned esti- mator for large-dimensional covariance matrices.

LatentLLM: Attention-Aware Joint Tensor Compression A well-conditioned esti- mator for large-dimensional covariance matrices

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.559064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.585546Z digest=sha256:6d532146c21d91d24bb33d9b77660feaf65f749031aacb9911fb65879dca99f1

Observation 583760f6-b70a-42c5-9c3a-8afa8482b7b6 · outbound

This paper cites LoSparse: Structured com- pression of large language models based on low-rank and sparse approximation.

LatentLLM: Attention-Aware Joint Tensor Compression LoSparse: Structured com- pression of large language models based on low-rank and sparse approximation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.384050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.683173Z digest=sha256:481a4796e0633eec24a9dcd11317e65b4e4a7390e16cf8cf1eb7d44b0e9cdd7e

Observation a6740093-4468-4bfb-98e9-ea5c20b7031a · outbound

This paper cites Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix.

LatentLLM: Attention-Aware Joint Tensor Compression Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:36:29.038387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.765346Z digest=sha256:088970d88a2c00c6b6e4b7d3f61c04f95caa9b3fe1ddc4ef78280c5b0cd91abb

Observation 4ebe3966-64f4-492c-a488-b3e3cf0c3119 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:26.860447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:26.860447Z digest=sha256:7921e69c9256a91df28a32fa0b48be89e9c7beef9784f0eb996ac44de7de8ade

Observation a7ad3544-b528-4917-b96a-df0dcc96819b · outbound

This paper cites AWQ: Activation-aware weight quantization for on-device LLM compression and accelera- tion.

LatentLLM: Attention-Aware Joint Tensor Compression AWQ: Activation-aware weight quantization for on-device LLM compression and accelera- tion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.213947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:26.946648Z digest=sha256:3012dc2c3b33c5925c9d4af6f250eebbc48e5de372ee7eda450b8f1863ea585a

Observation 8284d0aa-2043-4fd3-a339-1ebdb6dd188f · outbound

This paper cites DeepSeek-V3 Technical Report.

LatentLLM: Attention-Aware Joint Tensor Compression DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:27.024844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:27.024844Z digest=sha256:746eed48b65dcd2fa885ab89488b39f68bd212c7ec96d49feb7c89f178658a94

Observation e634fac1-8581-4ad3-b278-992e3eb03c67 · outbound

This paper cites Visual instruction tuning, 2023.

LatentLLM: Attention-Aware Joint Tensor Compression Visual instruction tuning, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:31.089171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.087720Z digest=sha256:a5ea3c397e4e8d6c4cf4f3551b9de4e5592a34a23c76d3b080ad313ef55dc7bc

Observation d47d09b3-8be8-4d51-87c8-b0a01a197364 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LatentLLM: Attention-Aware Joint Tensor Compression Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.921460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.150867Z digest=sha256:b7ee1ddc873df4c8562f068746eaf2773ef013e55c90e08adb8c1c3f4827a02e

Observation 932a920d-14f5-4e6d-9c14-3e22c81754f0 · outbound

This paper cites The penn treebank: Annotating pred- icate argument structure.

LatentLLM: Attention-Aware Joint Tensor Compression The penn treebank: Annotating pred- icate argument structure

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.811173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.221411Z digest=sha256:b7af65ff4e45623652926e41a4c3fc03e13f3faa505ad94403f9133517cf9bc3

Observation cf4e8bf6-d8af-4ab3-bd6f-b2f5d8778f46 · outbound

This paper cites Pointer Sentinel Mixture Models.

LatentLLM: Attention-Aware Joint Tensor Compression Pointer Sentinel Mixture Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:27.293105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:27.293105Z digest=sha256:bd08a59564555f0b57fec117a36f073300941f253749bfd70b875826b162e3fb

Observation 47530378-0076-485c-a6a3-5bafe57adec1 · outbound

This paper cites PyTorch: An imperative style, high-performance deep learning li- brary.

LatentLLM: Attention-Aware Joint Tensor Compression PyTorch: An imperative style, high-performance deep learning li- brary

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.655344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.355659Z digest=sha256:db3494c851b369065191363118821f779c1cf0f15c219fddbfa9e4eed67a11d7

Observation e5a97581-de30-45d0-8164-2ca0c1fb69ae · outbound

This paper cites Improving language understanding by gener- ative pre-training.

LatentLLM: Attention-Aware Joint Tensor Compression Improving language understanding by gener- ative pre-training

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.541603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.437977Z digest=sha256:7c2e5a8bc60843bb85aee5f90caf5b360b9dd0b69b0f0e1ae2a9fff32f96e8d7

Observation ded2fb4a-2627-4b35-b587-87a07647adc0 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

LatentLLM: Attention-Aware Joint Tensor Compression Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.403281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.503103Z digest=sha256:fb8a52b33c3c208e74432078165e7c871e3febba08732eb2437e5e484b573b33

Observation 308176fb-f749-456e-b4df-b5e993404494 · outbound

This paper cites Compressing large language models using low rank and low precision decomposition.

LatentLLM: Attention-Aware Joint Tensor Compression Compressing large language models using low rank and low precision decomposition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.218568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.600397Z digest=sha256:b49025c5cb9de8560db6f380c201d9111c974697ac07fc8c52166d4d5e000d11

Observation 81cb0c3b-46dc-41a3-88f7-f9d8c5d4fcb3 · outbound

This paper cites Low-rank matrix factorization for deep neural network training with high- dimensional output targets.

LatentLLM: Attention-Aware Joint Tensor Compression Low-rank matrix factorization for deep neural network training with high- dimensional output targets

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:30.045542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.683165Z digest=sha256:4e1b76ee80f01940b5adae06ad19f1343c926a8bc9ff4d434b18ee2d45272328

Observation 508cc0d9-098f-4f60-b8ab-01ccc9a87c36 · outbound

This paper cites Eigen Attention: Attention in Low-Rank Space for KV Cache Compression.

LatentLLM: Attention-Aware Joint Tensor Compression Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:27.748100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:27.748100Z digest=sha256:b73ba0e134702ebb8021505637a0932317cacd3d42865f9ccf173810651e1715

Observation 20dda85b-cc72-4e9e-9352-e82849ede1a5 · outbound

This paper cites Low-rank lottery tick- ets: finding efficient low-rank neural networks via matrix differential equations.

LatentLLM: Attention-Aware Joint Tensor Compression Low-rank lottery tick- ets: finding efficient low-rank neural networks via matrix differential equations

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:29.922562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.790671Z digest=sha256:d310d49fd148b5c255bc33380fafb95a5cb5154821ca2651625873fdf362bc77

Observation 2ec3fa6e-c35f-458c-951c-7b43c5e80af2 · outbound

This paper cites Green AI.

LatentLLM: Attention-Aware Joint Tensor Compression Green AI

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:29.800985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:27.854801Z digest=sha256:7845c8fac08ba5a1dd58e9f27ac9c4e7f314412334d50471ec2ce8d1b284d026

Observation 5e7a8c2e-c841-4268-b95a-cabab14488f8 · outbound

This paper cites RoFormer: Enhanced transformer with rotary position embedding.

LatentLLM: Attention-Aware Joint Tensor Compression RoFormer: Enhanced transformer with rotary position embedding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:27.942579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:27.942579Z digest=sha256:7af59ce6393ec3f62ab77fdc24231c25bb22b0fc5e9c633d82ce6bb85743ab33

Observation 18f6aca1-b490-41b1-aa38-f50ef6ff4c85 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression A Simple and Effective Pruning Approach for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:27.999333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:27.999333Z digest=sha256:265f0e3473c1d4331e63c2826de1c9ad4fa44007e8ebda3ffe3b60263bb6b096

Observation 4c43f400-fd57-4aa1-8161-ce727b13c736 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LatentLLM: Attention-Aware Joint Tensor Compression Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.069833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.069833Z digest=sha256:3083e2e119914a15724ccc7f75c7e01067e9b1360ae6241ca7411815ccc6c2b6

Observation 76d86c79-9ba3-471e-ae7f-c9e467fe44df · outbound

This paper cites Q-VLM: Post-training Quantization for Large Vision-Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression Q-VLM: Post-training Quantization for Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.109772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.109772Z digest=sha256:8e9a8a0ff68b602d4903784e62430b8776fad571bd06cf10cc2ba0aba031685c

Observation 603ecd6f-dbb6-4f96-bcc4-d59ffd567631 · outbound

This paper cites SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression.

LatentLLM: Attention-Aware Joint Tensor Compression SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.168443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.168443Z digest=sha256:0f0f4ba66d690fb0bb975becc3918d3e98fb0a97326ed43735b5a293d24d0102

Observation 7c9bb04e-07e6-4706-9bff-98ca4f8da6a1 · outbound

This paper cites Emergent Abilities of Large Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression Emergent Abilities of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.207771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.207771Z digest=sha256:9a6101467806c02b2a490cc172e40bf03d6fbb46bd8146148d7d01442c2dc2e2

Observation 34cbf179-46fa-4b85-84a3-6ce0b858f452 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

LatentLLM: Attention-Aware Joint Tensor Compression HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.266940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.266940Z digest=sha256:dab595c5acf19ad4cb325d7552c2eb70c6fec3434e62868a4cabe77b57261c65

Observation bb1f7871-a812-4f81-ac68-d52ca000203e · outbound

This paper cites A survey on model com- pression and acceleration for pretrained language models.

LatentLLM: Attention-Aware Joint Tensor Compression A survey on model com- pression and acceleration for pretrained language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:29.690027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:28.327746Z digest=sha256:af3d9d4a1deea881c8900be69eb9f2d46856ec37ee957d162cbf9b38e1d03744

Observation 657eba76-ddd8-4c71-9aa4-d7ed7b2be546 · outbound

This paper cites CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning.

LatentLLM: Attention-Aware Joint Tensor Compression CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.399010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.399010Z digest=sha256:4dc6b99e20a16330130aa3724adb6194358e86dd8320b5143b252f1a8b145cb2

Observation 9657a837-80cd-4eb0-91fc-86a062381182 · outbound

This paper cites ZeroQuant: Ef- ficient and affordable post-training quantization for large- scale transformers.

LatentLLM: Attention-Aware Joint Tensor Compression ZeroQuant: Ef- ficient and affordable post-training quantization for large- scale transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:29.560048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:28.478865Z digest=sha256:9f02dc011f7f64edc57536beb874fb4420b6f9d6577f19a91fac20b11526797b

Observation 58b88554-8138-461b-adcc-18dfbad1c2d7 · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.526259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.526259Z digest=sha256:c80f19fb671794fd28a5dfabc2132f4faadcb41e3e8bfeb94fb1bd81f8994294

Observation b93a57aa-0a94-4728-8bc8-8f5214d49e12 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

LatentLLM: Attention-Aware Joint Tensor Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.556122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.556122Z digest=sha256:554267b9508ae8ae8e4312077fbd8d773e8b9f513f977c1b393fe830ae6be869

Observation 26e09b5b-3bc5-445f-a80e-64b3cbf75517 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

LatentLLM: Attention-Aware Joint Tensor Compression OPT: Open Pre-trained Transformer Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.641006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.641006Z digest=sha256:6686c9a35b4baafa4d712d35e4599a519f24b77a552c62a2751e91a06e87ae01

Observation 242077bb-9032-41fe-a6f4-4057b8020261 · outbound

This paper cites C 1 2 O µ⊤C −1 2 (1 − µ⊤C +µ) 1 2 #.

LatentLLM: Attention-Aware Joint Tensor Compression C 1 2 O µ⊤C −1 2 (1 − µ⊤C +µ) 1 2 #

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:29.428122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:28.682828Z digest=sha256:d95695fff2edfc13de82b378930c5e6b0688b73112127d5f2ceb1e3fccb2a256

Observation 235a3222-a37a-4581-91b0-6192fddafa9c · outbound

This paper cites (193) Plugging into the loss gives: L = X i ∥Wo,iWv,i(X − µ1⊤) − ˆWo,i ˆWv,i(X − µ1⊤)∥2 (194) = X i ∥ Wo,iWv,i| {z } Gi∈Rd×d C 1 2 0 − Bo Ao,iBv,i| {z } Hi∈Rro ×rv AvC 1 2 0 ∥2.

LatentLLM: Attention-Aware Joint Tensor Compression (193) Plugging into the loss gives: L = X i ∥Wo,iWv,i(X − µ1⊤) − ˆWo,i ˆWv,i(X − µ1⊤)∥2 (194) = X i ∥ Wo,iWv,i| {z } Gi∈Rd×d C 1 2 0 − Bo Ao,iBv,i| {z } Hi∈Rro ×rv AvC 1 2 0 ∥2

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:36:29.249964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:36:28.735305Z digest=sha256:8f7546445e1e3ed82201cb1fbd0174a6d68f4489a0c345d97574a4569108f012

Pith citing papers

No inbound Pith citation observations are available.