Pith. sign in

Paper Citation Record · LEDGER

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

As of 20 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 4 inbound Pith citation observations for arXiv:2501.15021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15021 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:02.691635Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:35:10.578399Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T19:58:59.018078Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2907b0fc-0182-4702-b433-c119f298753e · outbound

This paper cites A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:03.010577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.593012Z digest=sha256:ea7f48cb3f08042277b9723502cf6c151e6dd3e63a6f85a4e8afcd04d16b950e

Observation aa69e52e-4501-4857-a2d6-2ed43cc9a2b7 · outbound

This paper cites Multimodal intelligence: Representation learning, information fusion, and applica- tions,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Multimodal intelligence: Representation learning, information fusion, and applica- tions,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:02.999843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.596641Z digest=sha256:6181bca20c6677d7a6744517dc378572debd87fd3ef6aa01a23b26a35f67895e

Observation f5a8b587-f08c-47c2-bf1e-b4f8f583f375 · outbound

This paper cites A Survey on Multimodal Large Language Models.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models A Survey on Multimodal Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.599980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.599980Z digest=sha256:940dd534a5849b91dea1c518e1f6bc4aa6f1b65806b7bbe99c2d0f8fac4c77ae

Observation 36c7e950-bdf8-4211-860b-43f7ecd34277 · outbound

This paper cites Vision- language models for vision tasks: A survey,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Vision- language models for vision tasks: A survey,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:02.989364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.603961Z digest=sha256:9f950c5c06be69ce197b4d2de007edd571c719685e566dd39bf0b7266a1df29f

Observation e5535967-32e9-4051-8d77-8d402f89e2af · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.607231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.607231Z digest=sha256:3fef3a652e19891c05f5883230017d43a0412155af796d2ecfc165a200438604

Observation 34a4dc67-f6ff-43a0-9299-eaaa39fe5e07 · outbound

This paper cites Attention is all you need,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Attention is all you need,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.610987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.610987Z digest=sha256:e160c38394c5ad70c5a78ee6bf2b171d61812774cc329529e058e64f9baf18af

Observation a579d485-6400-495b-bae2-795a7872a962 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.614371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.614371Z digest=sha256:2e9aae2fffa4e8502e42c076d6b0ef5b0827f8b36e9dc06a21630a799ab3a2a5

Observation 7365c393-920c-4980-8ba4-da75b8ee00b6 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.617582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.617582Z digest=sha256:fe95353e4a331db83df0d582a0692b43d788b92771f05d3b7a80919323c54e89

Observation 313df24d-61b0-4af9-bc3a-58f2b96edaef · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.620758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.620758Z digest=sha256:441858519e245b1aa452be12988e2945445ee354dd4670e9d4ce19d438b11f0b

Observation bc4dbbfe-97da-4ac6-ace5-5beb16544490 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.623746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.623746Z digest=sha256:ceba83dab1d3102846462d1094dd755d951b7fb5541602e81f1eeaffe5f31c5c

Observation 941b6830-364a-4660-bf14-9df404ed0d77 · outbound

This paper cites Visual instruction tuning,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Visual instruction tuning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:02.973827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.626805Z digest=sha256:5e6899438f12b8bebefcb3646293cce807cb9a7df773c67e4f7024fc81b28181

Observation 422ee90a-9303-45ac-9d62-6c3b679dc066 · outbound

This paper cites IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.630189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.630189Z digest=sha256:e01d9934a80a9d748d0d8d32e03862e1f38a8ca4094230eca92cd8c9fdf5cd52

Observation c8ef2a7d-6b2a-4539-a203-bfa076ed1026 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Efficient Streaming Language Models with Attention Sinks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.633146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.633146Z digest=sha256:a4577ed53071d634064facbc8f50caeffd464388c78de88732719dd8049f5215

Observation e98dbfdb-7b16-4efd-8aca-1176e80b66db · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.636154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.636154Z digest=sha256:59534476df1b24b4d485e43561f10c4e1362148a159ce78a3d6b7e194fd67e00

Observation 44d2c809-e1f0-4945-88ee-dfa0debcab89 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.639513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.639513Z digest=sha256:c1021d74676a91be71c0144a7b3310b1d12e1c9983de1914248162f62459475d

Observation aff4718b-c0a3-4b33-bcaf-1f30955985b8 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.642735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.642735Z digest=sha256:8b8ab0be6430f78238c96f839b4f3f118b2361dc313c9f10a3d619940f2c320c

Observation 9d01b6e7-0e06-4a96-9212-907b78e2c062 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io- awareness,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Flashattention: Fast and memory-efficient exact attention with io- awareness,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:02.964505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.646463Z digest=sha256:7e6fd392174e15cffa2b2ad676d0a7b28567445394f0b74415f0df8490759fd9

Observation 1683751b-9abd-411c-ac42-b2a68008c86d · outbound

This paper cites Pointer Sentinel Mixture Models.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Pointer Sentinel Mixture Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.649835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.649835Z digest=sha256:1ce34b4964bdf4d93133cb89f93c401d0fed4a9703b6ca6259f5be25c504df0d

Observation e23be23a-fca7-4ccf-9384-7ed714e1cc6e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.653368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.653368Z digest=sha256:f4934f8768a45151d4f9840c3371780461a0c47a0e6eda4c3688e824a3cec531

Observation 80883f06-b8de-4bdc-b98e-2f219fa9b5f7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.656855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.656855Z digest=sha256:8b740fdbbd774319a61ed51041c7a23648835505e859c4c4e0bc8114a7b005db

Observation 4e8d4ff5-4a2e-47c9-96c1-6650b7141b24 · outbound

This paper cites Improved baselines with visual instruction tuning,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Improved baselines with visual instruction tuning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.659900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.659900Z digest=sha256:746a9353987f82b6b517718318e594b311564c7a0b1685746bd426253a4d3b86

Observation 73b22388-3389-48dc-a8ef-d364a7654621 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.662595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.662595Z digest=sha256:b650126882713e6f4b13e51f895a2044fb69d5db5eebe9c75410d1905a9e3eab

Observation bcbb5214-0551-4032-bdb9-eb7d2ad6792c · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.665566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.665566Z digest=sha256:a54928bbab3454661f660c9c31e8ac11210d3d834a6fb53d2976db65adb65dbe

Observation 036eb6bc-19f3-4c1e-879b-2877d18b95b1 · outbound

This paper cites Massive Activations in Large Language Models.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Massive Activations in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.668507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.668507Z digest=sha256:d9ac781c66437750c8eb96bdc3167ce815355d38e166711c61352cd356dfb3b6

Observation 2cff7cdb-d691-4be8-b6e2-7f2fcb76685b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.671624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.671624Z digest=sha256:e69741f17a0a9f5ad1492aaad484ea3f3d5614546f2f1a624c81c7a50db0386a

Observation 6f04b087-f4aa-4ad4-8392-fd41af31a49a · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.674659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.674659Z digest=sha256:09890d5ceb5d447d2a593ebc26b77924e3a64cac8c442f77039d12fa17c8d46f

Observation f7d14ab3-8f70-4bdb-92cb-756da1020f2f · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models SpinQuant: LLM quantization with learned rotations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.677581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.677581Z digest=sha256:fa1259a0df72ebfc5ebc3385503c92b0c016ebe527af7a720aa1659714086451

Observation 61d6b4fe-7c93-47bc-85c8-1258a6985e36 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Roformer: Enhanced transformer with rotary position embedding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.680490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.680490Z digest=sha256:7cc8d771cc2139ed1ab481e75cb52ac5f124ebb0a113d0dbe47ffd1f5ac4bf8a

Observation 573fe177-decd-4d7f-a325-eeef130edf6f · outbound

This paper cites Unified matrix treatment of the fast walsh-hadamard transform,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Unified matrix treatment of the fast walsh-hadamard transform,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:02.942757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.683249Z digest=sha256:cdf49784ab9db31f402c834d5601e775be5113716394d480dfff3161272a34e2

Observation 0fa639e4-067c-4126-9e24-8803931dd96e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.686037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.686037Z digest=sha256:24126215e1c3d851aee6cacc8049224c7262586f89178919fd26cc2f922fdca1

Observation e23df800-0085-451a-ba68-199ae662ff8b · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models MileBench: Benchmarking MLLMs in Long Context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.688706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.688706Z digest=sha256:2b3884c5a0cdaae62f82a04b7f92c7a15d230890b15c13b119dce85f862c7a9e

Observation aaab400c-d280-497f-830b-38469e01e32d · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:02.927106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T14:46:02.691635Z digest=sha256:144f7dd01f29f6c9e43621d6c262fbffe55c3d83b7cbec3185217b1773aff7e1

Pith citing papers

Observation 6a3f4b03-3f62-4b49-a654-dd4d3dab1e07 · inbound

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity cites this paper.

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:35:10.578399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:35:10.578399Z digest=sha256:91a95434c77f587f4901208473fd1205424c087bbc6f4badb567119dcf088a9b

Observation 75706047-98c0-465f-8482-4571787ff875 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.761408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:24981e7d21bc98d776edbde6ce5c1775ba70d7892668dedd00e90b94a8ab32ac

Observation 66c56521-712f-4369-a62d-1b327c02f325 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.018815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:060be1dc0c658bb32d5a0c58d4a7a2f1f6b1881a3d3b3c0e253e84b334e13503

Observation 349a6c55-f8c7-4b09-9350-1f1978497543 · inbound

KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy cites this paper.

KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:58:59.023556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T19:55:35.051834Z digest=sha256:8b207c3b200f23be5d42c5205b0dac35751c30aad36e85be2e5841e72dfc0e6a