Pith. sign in

Paper Citation Record · LEDGER

SqueezeLLM: Dense-and-Sparse Quantization

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2306.07629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07629 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:30.128860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.340082Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe5067cf-dab7-4166-aa91-a11026099e93 · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:49:33.850130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:6b7c68b3801dfe310d5755f64b9ba2c474c4ab854b009f94cb9c638a656be235

Observation 652a7ae9-34ed-4e07-bc1c-d8e964f9b301 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads SqueezeLLM: Dense-and-Sparse Quantization

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.341205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:a6171e4be6b2a9e12750cc8713cfe15978c07cf21c85391fee23f142b382e7c6

Observation 37c8d920-1703-41cf-9a2c-069a3b3898ab · inbound

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache cites this paper.

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:53:12.337562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T08:53:12.253243Z digest=sha256:3069304858ede98a813dfd12c5aea82e2a7a20027896d6f79b288f2f7ee99556

Observation 8c46ab81-0c82-4bf6-82d4-4f4f738c1f12 · inbound

RouterBench: A Benchmark for Multi-LLM Routing System cites this paper.

RouterBench: A Benchmark for Multi-LLM Routing System SqueezeLLM: Dense-and-Sparse Quantization

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:47:31.094875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T10:47:31.006944Z digest=sha256:81f5233443753c7568fa89a5f145969ec480c4ec70364e60dd3b05857eff9c33

Observation dc0a389b-ad0f-4564-b0d7-fec01552fe30 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.214075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:8fc6c9296ffaec410498caa5d01524ddeea6996af36d4027ac9ddfb7b6d2d703

Observation c2b6fba9-5454-4565-b948-54fe9cad1fec · inbound

SpinQuant: LLM quantization with learned rotations cites this paper.

SpinQuant: LLM quantization with learned rotations SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:52:34.700791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T15:52:34.606853Z digest=sha256:4d863aec99e27a91cc72c868cea3f88199bee34780fcc3e2c75448d82a99ac04

Observation 7d957495-3395-43b2-b49d-98a464d2b749 · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.128860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.128860Z digest=sha256:bab0b2178dfd857fd59e310541f455faabbaaf3c8cf47887d5ffaee6e869d221

Observation 546c6e91-2378-4d7f-b935-d5c1767416df · inbound

FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices cites this paper.

FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:20.735632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:56:20.735632Z digest=sha256:2fbf2c34db0630d3d9b12331c55de01010a84ff3c4ddd4b19cab000a738989e6

Observation b321e75c-9bc7-4ccf-b55a-f6121396b34b · inbound

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks cites this paper.

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks SqueezeLLM: Dense-and-Sparse Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:43.264913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:43.264913Z digest=sha256:fae1584046b091291155f9d8d942e99c49717a81e91e9c7371de00eb649bcde5

Observation e4589b22-bab1-4780-bff3-fd0e67da1f41 · inbound

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring cites this paper.

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring SqueezeLLM: Dense-and-Sparse Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:57.446387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:57.446387Z digest=sha256:9979440f61ab61171debdd4e24ecd6374694b9a256f9f3669d5af4f8fd9b6e13

Observation 8a3ad5eb-fb11-437b-a842-f2b5f1e3381a · inbound

Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale cites this paper.

Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale SqueezeLLM: Dense-and-Sparse Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T21:23:48.574835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:23:48.574835Z digest=sha256:54cf67ae27529693df3fcd3bcd004fe6ad7475ff39432060424ac3061443f0bd

Observation 8b308851-e505-4df8-ad27-7ad194f53a52 · inbound

Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives cites this paper.

Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives SqueezeLLM: Dense-and-Sparse Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:28:42.183775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:28:42.183775Z digest=sha256:83ce814601025ffcf877a17d18b6e12b8cb8933a4aaa2f36e34a6125770f6d1c

Observation 82638e70-4319-4cad-85ea-af7a65b3dd55 · inbound

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding cites this paper.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.010420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.010420Z digest=sha256:cfeb44a565b7f1db5aab706c3b0d35b591ce095ee012227573d519b7ac1382a4

Observation 82fa24fe-d6f5-4e80-a3b1-bf4e25e3196a · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters SqueezeLLM: Dense-and-Sparse Quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.544467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.544467Z digest=sha256:40d0a6ff1b42856db4aa198a41af5291b49c9f2bc2f2820cb46c2de89a65e254

Observation 4eee505b-e77f-4694-be24-ab55ae4e2279 · inbound

On multi-token prediction for efficient LLM inference cites this paper.

On multi-token prediction for efficient LLM inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.760608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.760608Z digest=sha256:9c1c7254f99966dc618a08bd7147b9d0cb290bf85e9a98e61204b280f0fd5522

Observation 9949b154-5c0b-4936-8e4c-06b655bb5173 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.938479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.938479Z digest=sha256:847c641b9829b103836bd91db7d338c4be8b45f9dddcbd2706907b9d2ba3df79

Observation a609dd30-4a14-47d9-b2e7-a04c77294da5 · inbound

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices cites this paper.

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices SqueezeLLM: Dense-and-Sparse Quantization

Reference 174

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T01:05:16.310998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T01:03:26.037233Z digest=sha256:a1b3fa53d03e2d9f20e2c1db09f8864c30e5bc316e95b845021ad9894dcae6ce

Observation e28853ba-9226-4caf-bf9b-35df3a4991a8 · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate SqueezeLLM: Dense-and-Sparse Quantization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.277661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:0ab361362bdf1d113429e36ef9c848758f3106330d78a54e0c3fe394713e6059

Observation 929dcb5e-aab5-4020-aca7-84de10b3c0bc · inbound

EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices cites this paper.

EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices SqueezeLLM: Dense-and-Sparse Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T16:34:59.261977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T16:34:37.083239Z digest=sha256:51c1ac0404ff5c738c48c51f5198a8bca85320e8b2778d0206db2cb52414b26a

Observation 776bbb39-0683-4ba8-a428-617b2fe31cc3 · inbound

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference cites this paper.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.656359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.656359Z digest=sha256:75a8b538634fe6a23ed020e5941010bcecde37e166c3a4b515a582735bdef1e1

Observation 61b27dfa-542c-4af4-b9dc-1325f222f8d7 · inbound

NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics cites this paper.

NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics SqueezeLLM: Dense-and-Sparse Quantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:28.498086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:28.498086Z digest=sha256:17e80bc33596b8764ec5eef96f08253742cad19388710759cab90315a0d25e1b

Observation c33b91b4-7076-4fed-839a-2d4ea3d92494 · inbound

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache cites this paper.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.069575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.069575Z digest=sha256:fda889c62b6c66e33cf41da0898e09fa4dab4dd56d15d8b08c613f36b5e2aa40

Observation 522070b8-6cf0-4e0d-a01a-26bae17ff451 · inbound

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression cites this paper.

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression SqueezeLLM: Dense-and-Sparse Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.372336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.372336Z digest=sha256:240f75c73ecf8b050f8ff56ffbb6a969c9e0841ccb49142c261bec8203ec3cf8

Observation d2d1cc3b-7958-4a42-afa3-9fa22aa93fd5 · inbound

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression cites this paper.

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.464514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.464514Z digest=sha256:a33ff1384ef1894a877147d2e73771b66a9e87f505f625f192d7daf92cfd72cb

Observation 8667e534-2518-4bf2-954d-47bdcf62d9c7 · inbound

FPTQuant: Function-Preserving Transforms for LLM Quantization cites this paper.

FPTQuant: Function-Preserving Transforms for LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.395454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.395454Z digest=sha256:02b91f3271d37d7e46f18fdc76690a123ba8d74d23c0a4097e6c2d5549d09acb

Observation 35aa94be-e2a4-4b81-aa83-0b824d01744a · inbound

BAQ: Efficient Bit Allocation Quantization for Large Language Models cites this paper.

BAQ: Efficient Bit Allocation Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.624716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.624716Z digest=sha256:9e5ddac3b7555196f8a7be157ad5a25cbc2257bc31c59c07373333dd8a7dca79

Observation 0c4f7916-3238-492b-bb86-7a95c6562580 · inbound

Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs cites this paper.

Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:53:04.144810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T08:52:45.818050Z digest=sha256:7901d94466e040ecf405b2ec3695315cae291837e19cc8e3c54c15b92e58f0ad

Observation 53dfb406-e291-4a4b-badf-f4e1a106241a · inbound

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models cites this paper.

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:19.558832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:19.558832Z digest=sha256:ce9fd5e37c562e8f9363b2874db5b19de1713ccc536cace8b9c49d4c36d6a43e

Observation deb9f2da-2284-4ec8-bf66-743303f6bb47 · inbound

Information-Bottleneck Driven Binary Neural Network for Change Detection cites this paper.

Information-Bottleneck Driven Binary Neural Network for Change Detection SqueezeLLM: Dense-and-Sparse Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.158945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:24.158945Z digest=sha256:59fac1f7198187c424890e5385f4f78cafa2d9c730c1393c89066f2ca940cf36

Observation 14813428-cb77-415d-9713-4f4d6f3f670b · inbound

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs cites this paper.

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:14.550295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:09:14.550295Z digest=sha256:98e2cfdeb176709327fb11c3e555b8b0e5a0d3475b00ba9d85b87a1f2798795e

Observation e84da4c0-1e3d-4d94-82c6-e8450d5c288a · inbound

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration cites this paper.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration SqueezeLLM: Dense-and-Sparse Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.956446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.956446Z digest=sha256:aa76e4bb418f88e40a3dd7db7c93bb8d865ee95fe9b82fb17035e10e807455d1

Observation 7aa8ef5c-c2e7-459f-b440-804064a16c33 · inbound

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs cites this paper.

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.950132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:52:44.950132Z digest=sha256:6ba782d29f227268a6f1681f1e663486e896c0c9d33d5440449e3686d3050f03

Observation 5345532e-07ac-4710-acc3-c0145d49c1b6 · inbound

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference cites this paper.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.877669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.877669Z digest=sha256:20215ecd99391aef2bf65e6129f2438b807d7d00159208950952cff4ec553eae

Observation 0985625e-76f9-437f-ad97-54c08b11d733 · inbound

Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation cites this paper.

Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation SqueezeLLM: Dense-and-Sparse Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:38:22.663076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:38:22.663076Z digest=sha256:cbac8f78f5b9fa0cf94f511644528ff770090f513c8bf0515a514f9c72993363

Observation bd74cafc-1369-4735-b1e1-8dc004cf81fc · inbound

CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization cites this paper.

CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:52:28.393534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:51:18.629467Z digest=sha256:b954bfa3a520eddf462261203f4246cd6dcb8cd799991eb70af67f006d16e0be

Observation 4f1281b2-6e40-4ab6-8b75-236e1bbd952f · inbound

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks cites this paper.

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.003271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:21:30.101748Z digest=sha256:5ce49791898196081db5e9985c415acc2c7eb8a19427a3fa9f68961bfbae3ee1

Observation 4de3b8e2-c80f-463d-a5d6-a86773e97d0d · inbound

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels cites this paper.

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:16.200257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:35:02.120009Z digest=sha256:8c278efd790b3d5af3284c6d7fdf25ba077b2b503b18e69f302e79294552cfd0

Observation e07df099-eae8-4bee-9a7e-d68669a9ec3a · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment SqueezeLLM: Dense-and-Sparse Quantization

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:42.443324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:175144ef213d7c10c4e4383ea19f588d5f9092455688acab3937d27284fbbbc5

Observation 8f50f0c4-d94e-433f-ab19-6021234e0e44 · inbound

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning cites this paper.

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:42.160072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:49:53.432959Z digest=sha256:e2723bff60004e329327cf0c44463750a4c35c022748e188fbc780bfa3abea4e

Observation 39e3d22d-1c77-4575-9115-5a4f1c462bc9 · inbound

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning cites this paper.

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:50:51.165324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:49:51.021741Z digest=sha256:1fc6ea2537b502efbfb1b04bdf111ba77deb0254349ecb4a919c62cf14888ee0

Observation 99e832e7-3ec7-4588-b1ac-415de378cc29 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SqueezeLLM: Dense-and-Sparse Quantization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.029916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:7662f2dba7bfff05e72f5218e04c680c6b85a0b49353493d0822c530fc39335f

Observation 0d0da963-e2f9-4554-b344-1f115737f3f5 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning SqueezeLLM: Dense-and-Sparse Quantization

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:10.321373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:4f71dd9d1c85c4ef4cc5c79dacbb0031bcc1f0dc71f988d6759682adf27cc08c

Observation ee2ac829-f9e2-4fd6-ad98-c96f3e2bc34f · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:30:44.142303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:23:14.935801Z digest=sha256:968680c2eac2b703e32e6d35c8ec6b824fbdce5d4c4a35f3ed98e864f10dafec

Observation 2740b635-abd5-496a-a944-9bdd0dd5b5ea · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.210120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:59:00.997742Z digest=sha256:b605bee80352d6411ef2e0bd13deb7a028eded2d5fc6a1515849bfbd970e249d

Observation d16da3e7-de78-4192-8548-66862fc2ff1b · inbound

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference cites this paper.

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.690279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T21:28:36.358474Z digest=sha256:ae56d0aa7226fbc8a2b5833c678b38ed401c1cf3d1c3081770ebb63084d99aa6

Observation 7ec28d36-0207-4a96-98b0-9217129dccff · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:23:03.625863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:39a966bfee3958c66cbc24e4f370787f26f383de732a4cdbb856eb67198d17d7

Observation 9bd3ab0c-6349-4222-a3c0-7f1dfd3eba6d · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.872360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:c2a2d5f8399bacfbecbe84bb530d4e310ca53022787b96718333650acdd01a89

Observation bf0cc52c-1c31-4da3-9a8e-9cba87bf348f · inbound

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models cites this paper.

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:04:58.491515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:54:56.386488Z digest=sha256:f0bf20a6df312be912dace1f80a5454e3b59a2725ba309e053df953006424ff4

Observation 9f262736-e8cc-4a2d-ada2-88c261a3f7b2 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation SqueezeLLM: Dense-and-Sparse Quantization

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.617720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:b8df5fa2d6f47e3978d7731caffe90add1184190d28518240307f18c15c2bbd5

Observation d4b4ba6f-6db8-4589-807f-f054a40c3feb · inbound

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation cites this paper.

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.475000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:28:14.160341Z digest=sha256:351e977c637af6bfbd7481be5cf3e92259a016d16ff95e6dc6dc1e7c8530a468

Observation cb5672a2-5a67-43de-ac03-95457cefd64a · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.293704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:08bda3bd02548ec9803a777643aa0618cb76ad256611fe84897f7ffcdf797064

Observation cc0d0cdb-359f-4fbf-a642-c00e4281e965 · inbound

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs cites this paper.

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:25:48.408417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T01:29:42.919461Z digest=sha256:b733dd9f2668d69dcc5dd17536f992a5c867ddd7af7e5aca1d25acb82b7c4888

Observation 1bac6c64-6ee8-4ee4-a0bd-310a5e10279c · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.403971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:9ee6549b9cf059068dc519a982e33ba8b0510db8efb50650d37dd51fd8b12cb5

Observation 807090fd-1b0e-484d-b714-c291eb299548 · inbound

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models cites this paper.

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:18:37.341488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T16:14:03.717787Z digest=sha256:9ca7bb37842758cc1e271bc4e1baa826a62f07d507e5623dcb0bad3181ad46dc

Observation 0e56bdc7-8bc2-4799-9fed-944eb3b39dbb · inbound

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers cites this paper.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers SqueezeLLM: Dense-and-Sparse Quantization

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:32.417750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T14:56:10.553212Z digest=sha256:6880ebab17795be090f0f67f3b0069a80aa53eab24155843222e55f551902350

Observation 91ab3514-5ab9-417b-bd6b-a58cc4aea220 · inbound

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration cites this paper.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration SqueezeLLM: Dense-and-Sparse Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:60a811aba81f90def58fe3e08169c95c32267504b3d8452ccd4a9078f8194e61

Observation 25b40a43-1701-4822-b146-c3fd2fdc511c · inbound

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference cites this paper.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.941291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.941291Z digest=sha256:478104e3357d7d1992186a7f99975ce89ef70437879e9a03e82f7eccd20276b5

Observation a32228b5-c235-4152-960b-a3ebecdb9bf7 · inbound

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications cites this paper.

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:14.348367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:04:14.348367Z digest=sha256:6b889b1da9b27f38d5c961e592af78db6b413c5c6fb4f6872f4f6795cf79882e

Observation c46894f7-5422-494f-8ba4-abcadcbcafc3 · inbound

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs cites this paper.

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:44.530504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:44.530504Z digest=sha256:87f6a6025a66ab9cb990361cffa0f5e5031ae7d52cdf84c3c37bad0852d8ae05