Pith. sign in

Paper Citation Record · LEDGER

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

As of 14 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.18231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18231 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:42:23.114204Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04021691-f782-4ee0-911a-c5d942775906 · outbound

This paper cites GPT-4 Technical Report.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.441861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.441861Z digest=sha256:2be09b0b3568f64786b2067d61801bcd35d0a3c7790e37b4b28b71440662e9f8

Observation c064fde5-75fc-4a2e-b708-2129e03f7fad · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.876599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:19.520380Z digest=sha256:0f574801a6c05dd57517221a3da9d599faad70512e61020288135ffb5496f48a

Observation d1864e14-059f-4226-8a0c-6a126f8d17bb · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache LongBench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.751190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:19.618799Z digest=sha256:d3c1eb9229a60cf99cbbc4fe0eb0ba01b6587bc1028c78b19a867f8bb148638f

Observation 88c24244-3562-4128-9757-43225bc40b2e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.668535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.668535Z digest=sha256:b2fb601b2aec758053b4c0789cfae8673472fcef86c2d5a7626118a7eff59846

Observation ba5f5ffe-9949-4414-a9f7-cecf98a97fd4 · outbound

This paper cites Palu: Kv- cache compression with low-rank projection.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Palu: Kv- cache compression with low-rank projection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.582959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:19.709613Z digest=sha256:ca4bd6ce6747a3bc2595b65726da2b32d1eceae27b58814b0daca93c7a6c81cd

Observation 91388cae-ec1a-4f50-b9c5-1d2c76dc15a4 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.764346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.764346Z digest=sha256:98dc76dc3a5682a5ce6d4f35a2b86e0991743edeaba8eeb5dde3e9bf25f0db39

Observation be3bddb7-379b-45c4-ad85-3dab61ece3b6 · outbound

This paper cites PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.813220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.813220Z digest=sha256:f4c57255894e0445dfa876dfe2fa8dd3dc7ea53bfbf1667dfb92df4fceffb6d0

Observation 589311a8-858a-4cf4-86fd-0d9430dc88b7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:19.966304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:19.966304Z digest=sha256:2acd4375c9da885ef5bb8cd9852a2fb93033d2b52f6d8ac226f7436165533e4d

Observation 523759e0-e41f-47d2-b35d-306f8731b9da · outbound

This paper cites SDR: Efficient Neural Re-ranking using Succinct Document Representation.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SDR: Efficient Neural Re-ranking using Succinct Document Representation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:42:23.794083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:20.113131Z digest=sha256:624744bc7684ef865a81971e923175f8212f73d022841d5373b7da963d1ad4be

Observation e6527f5f-df86-49c4-b3d3-88fdb0cc812f · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QLoRA: Efficient Finetuning of Quantized LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.170685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.170685Z digest=sha256:2aaf33e72f28d88b1c3e2400ef8c682cdecff224fcb673aa8abb77d0031f5080

Observation daa42476-cae2-4cc7-a772-afad805be48b · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Extreme Compression of Large Language Models via Additive Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.242101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.242101Z digest=sha256:807e64f7e14a8342ee0ec384bc25a9811371655cb46ab3ca598b68c22ba4375a

Observation 20dde1d7-2727-4da2-8f76-345abae103fd · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.321754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.321754Z digest=sha256:5149efe97795bc97f98d1a97ae68ee141266886ef1100e3ecefbf26ee2ba1b6c

Observation 7b40e5c0-9aa5-4cf1-b909-8b9a77f3007a · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache A framework for few-shot language model evaluation, 07 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.414042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:20.403085Z digest=sha256:fdd4e0efa0b6d67fee92fd99d5b988cd36a204a10e8fb0bba20d608544b4f531

Observation 3b8af31b-88bf-48d7-8a27-2733d47331b5 · outbound

This paper cites The Llama 3 Herd of Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.499175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.499175Z digest=sha256:50852fee54ce07b170c157444f8ad4c5d1f5f5e79b74f9c9c56efc3923905128

Observation 0e792160-6fcd-460f-ab99-611bae593cf0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.623145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.623145Z digest=sha256:1a24bdbf1b54ed82b4381d13e2ac930829fed6185e147e9c1619adcbc26db5a5

Observation 96eb77fc-124f-4476-92df-21cf9a40f843 · outbound

This paper cites Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions.SIAM review, 53(2):217–288, 2011

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.725384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.725384Z digest=sha256:66772a1ef5a9787a48e58e872cde7d6bb67f822873e12555e991fca09b5134d4

Observation 735a74b6-9c5a-4958-af5b-36956bc96a53 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.783045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.783045Z digest=sha256:1be736b11f564b92363a610c7f16f897f77fa5357f66f29505024bf33eebc260

Observation 4a3d52a4-87ed-4ce3-9d03-863457c7e3e0 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Measuring Massive Multitask Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.854098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.854098Z digest=sha256:778662dcc714c8e1045c3f0958522ce3c9887ab292a86307ddbc0298068e2044

Observation 5ca32f2e-10e1-42c3-a799-bdc70c0a2837 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.904104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.904104Z digest=sha256:c4203ae96dc74b8a00362fd962c53f4b9c5f505eb5a064b7cd2ce79cdb13484c

Observation 67e9cea5-60b0-4f83-a1c7-19a6a6315861 · outbound

This paper cites OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:20.965959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:20.965959Z digest=sha256:01c0c9fd7f0b649e486abcaec27574d554dcd01324981e0fd0cf88adbaa8415d

Observation 0ffa03f2-aeb4-406d-82c3-eb69fd9ac9d4 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.019251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.019251Z digest=sha256:e658c2992cf5fe5e122fc87d942261ba96ae8134508cc174b34c21489c5ebe4b

Observation c33b91b4-7076-4fed-839a-2d4ea3d92494 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.069575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.069575Z digest=sha256:c7b8f962fdbcd77bc354a8a659036f5e1b0b073d5a063ba3a17a91879ff239b8

Observation 510a1e5d-fb85-4824-9e6c-b785a122dce8 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.150721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.150721Z digest=sha256:7ba8579a8af53d52e174533e0b002d052d061604f3b850478fedcab3d116be43

Observation 93600400-dce6-4a09-8da9-685a63ff55d3 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.194134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.194134Z digest=sha256:3d22a43c06eb0facc298d5c5e7b509ae0d43527061267289aa4f3233ac4eb5ed

Observation 1d7e3402-796f-45cc-8aba-8a6650c27255 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.275247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.275247Z digest=sha256:79672dd92dcb13c47460aa9751e07e6b551f2d11492ac21471d5941eae6b39f8

Observation e30a4de3-986a-4d57-bf11-6c0d9001326c · outbound

This paper cites MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.377519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.377519Z digest=sha256:e7cdff2eb1aaf33f2878dd2c4aaef62c30a7e7414534008db7c81891f6296317

Observation 7ff29666-6aef-41c8-840f-82f242a95a2f · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:26.198825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:21.476258Z digest=sha256:a1d4206dc43aa483ea671c389e340b4a5f57bbf48403eef5f6c83cc305507791

Observation 0c0a7fbb-d6e8-41a6-9cc4-bf018e610764 · outbound

This paper cites VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.528632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.528632Z digest=sha256:f8baf10a0c226e613a7c84be3482c93716b77447c4beaf0799885c9318d3d9ae

Observation 0d3692ca-f0bf-41ec-ae6a-a95c3ae19497 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SpinQuant: LLM quantization with learned rotations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.563050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.563050Z digest=sha256:aa7542e7cbac1b8cbb3ada098b89557cc4e782042fa3b11c4c39e0bf859d52fe

Observation 86282d67-e22d-4b85-a153-824487fbbdd5 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.631333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.631333Z digest=sha256:673ac959cb72eea95f986f4412fbb202bd51f75d9f8ae4e8be7d37f6d97a1719

Observation ef42813c-e02f-4efe-a7e1-e6948ece603f · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:25.985348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:21.706702Z digest=sha256:b4e5f3b15cbec2e68258b712f0b89ad9983f1a3218b1eeea268e5780d6160eca

Observation 8a2d3072-513e-4c5a-8db2-9ed89116cbc4 · outbound

This paper cites Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Cake: Cascading and adaptive kv cache eviction with layer preferences.arXiv preprint arXiv:2503.12491, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.756963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.756963Z digest=sha256:0ebe6501f81b2e2bd97dfa9e1b2232a13d939accf73a02e86833f5fee235c24d

Observation 4c80056c-736e-482b-884d-e635ae7bda2a · outbound

This paper cites Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Coqa: A conversational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.846993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.846993Z digest=sha256:dec74e90b6913d0cd659ca9b31565c023521fa38714f29ba5bf51d1d9bd58de1

Observation 93d8bf7c-4959-4bc8-8d67-58a6147bd157 · outbound

This paper cites Hanson-wright inequality and sub-gaussian concentra- tion.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Hanson-wright inequality and sub-gaussian concentra- tion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.738176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:21.943014Z digest=sha256:60c2aab6aec4d3dce790369b60af73a69a5dc8301c95d98926cf134b0260b128

Observation 8a99de4a-d23d-4050-850c-d43e52fedb8c · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:25.616849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.054272Z digest=sha256:48d1cced8d1b4ea75ea206d42c71ac7071c4651b91f9f7b99092c1a0bdbb6c41

Observation 9e15ebc8-9549-4bb6-a683-c712d4cb7235 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.125724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.125724Z digest=sha256:f783dd00137193a15bd17444de6577d3a17f19fe665ea320e17a4d470f3b1deb

Observation b0a527a4-b230-468f-bd55-1992104ff61b · outbound

This paper cites QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache QuIP#: Even better LLM quantization with hadamard incoherence and lattice codebooks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.408508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.217465Z digest=sha256:034d8eaf924660093e32976ebf1cc24691f516017c5c1395c2a18042d2239f0a

Observation 07b96d21-7971-43bc-ac17-0773479f1a63 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.302830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.302830Z digest=sha256:150a57ae59c9b328004daafda35a467201290aeb3a183c55835fbd316cb09ece

Observation 50e1376b-fe23-4b48-b5f8-6a7a9f9b03e5 · outbound

This paper cites BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.378319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.378319Z digest=sha256:7a40559f9a1b8770064b94a217134ab4db6f65746c165eaced75ceb40f089582

Observation 487e206a-3751-47e3-9ec8-a14135340265 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Chain-of-thought prompting elicits reasoning in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.452938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.452938Z digest=sha256:7faaeaff8a6b6befb16adb2f9913a7f65cc331a37774c656c7f5594ed19ec50f

Observation d96754ce-5019-4912-aa00-8454d5ad2387 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:22.512307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:22.512307Z digest=sha256:131900120f6aa6b8269987354afe4aa1579b84ee6ba80dfe19425743fa23d937

Observation 87d1acc8-019d-4a99-bedb-7d94516415f0 · outbound

This paper cites Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:42:23.446611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.560543Z digest=sha256:55c3eb0cb82d4c332944b189e495666dd4ea34026e331b3b0ca991e1f60c1758

Observation 96f64d5c-4832-4640-93ca-2772e53b30f5 · outbound

This paper cites Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Kv cache is 1 bit per channel: Efficient large language model inference with coupled quantization.Advances in Neural Information Processing Systems, 37:3304–3331, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.245202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.657980Z digest=sha256:5389b856c266e74a7bd49fe3549db348e93f3120baafc8e46d133ae1eedf70bb

Observation b3f8a2f2-0842-4cc6-93e4-361c804c15e6 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:25.047251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.696914Z digest=sha256:f44b8f3efc036cb49d957c1272e1bc46a7e70097513677f9ad7464b795227701

Observation 4c769c61-a318-4598-9404-46aaa563c585 · outbound

This paper cites Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Decompose Σ = Cov(X) =D+A , D= diag(Σ),A ii = 0,∥A∥ F ≤Γ

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.860495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.744073Z digest=sha256:537fe054cbeeebd41ae6eeb02a9dec53ff2f76414ef67fd8b59d2718a298acee

Observation 005a9e7c-b858-409e-88b6-4d08833ba59f · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:42:24.627060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.835376Z digest=sha256:cb93d49cd2eea2e872e12e4f98e3da52d28da96d2ccebdbede7062301308be66

Observation 9ae2d4ba-a8d4-4bcd-a9df-c8af1065cd99 · outbound

This paper cites Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h).

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Then hTDh= 1−¯ε, f(h) :=h TAh= 1 d sTAs, and Var(Yi) = (1−¯ε) +f(h)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.481106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.891916Z digest=sha256:161ae06d1267f73d5a2c5adcb8f44f515798c609a37600a3379a67e79409775c

Observation 3ff43d2c-dac1-43d0-9560-f3cb61434e5a · outbound

This paper cites Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Applying the Hanson-Wright inequality [34], for anyu >0 Pr |sTAs|> u ≤2 exp −c u2/Γ2 wherecis a universal constant

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.295084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:22.978513Z digest=sha256:cac553329c2b313190ac0525df1512115138e8aae7968268c5d11a104328b5f5

Observation 94c89684-f3fd-4a5a-9186-136648659cb3 · outbound

This paper cites Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Exponent becomes −ln(2/α) ; hence Pr(|f(h)|> t)≤α

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:42:24.086439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:23.036088Z digest=sha256:4f336a0d6c4f13014bc503c6adbd2ca44fed9beafda029d99295c6fa85d02ad5

Observation 9145d042-9aea-4cae-9d61-0018c71de493 · outbound

This paper cites an unresolved cited work.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache Unresolved cited work

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:42:23.342602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:42:23.114204Z digest=sha256:9f68aad344dc9af7221bafb1b95b4564823462e65502f53595c73346ee1ccb6b

Pith citing papers

No inbound Pith citation observations are available.