Pith. sign in

Paper Citation Record · LEDGER

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 9 inbound Pith citation observations for arXiv:2412.14363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14363 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:22:39.003914Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:48.606027Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:52.356456Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fda058dc-91e5-4449-8546-40e713c32df8 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.755084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.755084Z digest=sha256:8a7550fbe55a046e14fedf2b97e6a98fc50ce7fa9c0128231695d13e88dc1ff2

Observation 98d6d846-a6bf-4649-8968-14bffd0b998c · outbound

This paper cites QUIK : Towards end-to-end 4-bit inference on generative large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QUIK : Towards end-to-end 4-bit inference on generative large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.760201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.760201Z digest=sha256:30f65eea85fb70100fad432c2fd375c82b5c14672b0d9ee3f7a184ec04c5e4c6

Observation 29ebbb2a-601c-4599-8f5c-79cef0d22026 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.763953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.763953Z digest=sha256:54639695063404b069a60da3df49f6a0eb2189ebe099fdb7959ccd3a71706608

Observation 0cf70a50-b769-449d-a644-eacf1bbe934c · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.768700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.768700Z digest=sha256:a3b04dcf0522ab727ec575808d80f15fe23da8f8af3328dddb5d493b5cc6d1c1

Observation 4906daad-b726-4035-baa8-4ed936460af0 · outbound

This paper cites PIQA : Reasoning about physical commonsense in natural language.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals PIQA : Reasoning about physical commonsense in natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.910642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.772381Z digest=sha256:368c71b73eaf6c45b21a644313369e357100015470002775821a0f5c6f6a5565

Observation 323c0ab9-4f1b-4b74-ad4e-9bb33b0d543a · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.776013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.776013Z digest=sha256:820390d9b1e03767dccdafcd4cd6c65e3634ec1467a9ce646c11a98418ceed80

Observation 73f4de28-cf69-4dbd-9b85-e5de4284169e · outbound

This paper cites an unresolved cited work.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:22:39.899036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.780084Z digest=sha256:60e1d19aa63ef6aaad26515a84e2e5c75d81f119197ee6c87a0a2220b30b5e83

Observation ea7bf735-f1ff-4498-b177-db544f4e548b · outbound

This paper cites PACT: Parameterized Clipping Activation for Quantized Neural Networks.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals PACT: Parameterized Clipping Activation for Quantized Neural Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.783535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.783535Z digest=sha256:8e34cfa04d2750ed56c545301e2d89d571e370b5cce893143652501b8b8d84b1

Observation b5a31317-a4a4-42c1-a1c2-c64253763a13 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.787784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.787784Z digest=sha256:058d7c5336d7bead03acf174f3be9081e28d12162556d0165ce4f1db86c53b55

Observation 888b947e-c520-44d0-b28f-42183e4f8769 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.792466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.792466Z digest=sha256:7d11a1be5c7d3fc577226d6c0d498015860bce3ac4dce617884cb4882b70578a

Observation ca9a8a28-4d9b-4d4c-9079-ec3c659ea1e0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.796468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.796468Z digest=sha256:0e9ca952b05247ab60a6dbd1ecdda78800e11e0f2cb6378722107ede8bc2d838

Observation 9735236a-e432-4f9c-a318-640e13cfae42 · outbound

This paper cites an unresolved cited work.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.800077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.800077Z digest=sha256:e49c33da289ea6a07bcc656d15db14d865dc323fbcaf85f256133d7fa1bdcd0b

Observation fa5ce59d-3650-4b85-9f2b-1675639fee9f · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.803446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.803446Z digest=sha256:3221dc4a2bca3d3e2f5db5d5751255d96987510149798a9af537b33101be1130

Observation bd55bf82-363b-46c6-ba82-1cd0ece397af · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.806973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.806973Z digest=sha256:afb31328c8ba746ebbc7107d5ce004699962ca17c8d833dd45b3ceaabc11748f

Observation 012ecf26-b6d9-431e-96fb-17a1f35b1b75 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Extreme Compression of Large Language Models via Additive Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.810656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.810656Z digest=sha256:8a4c4d741322c72f24cc3248d74c1b7c119be839bc48a1fe9412eea4cdb03817

Observation 0300c697-a7cb-4e93-808d-7971eda845a7 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.814251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.814251Z digest=sha256:0110cb5e3c5149dba7682e1967fd38b5175f484c50c1e04b38563292fd5e7d09

Observation 9320ed35-7913-479d-8256-8bc779be1619 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals A framework for few-shot language model evaluation, 07 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.817747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.817747Z digest=sha256:47cc3888c99254129ef80d71ce153d40709cba33729e4a846d87e6ee1240441a

Observation 98a1aaa1-5cee-4e4c-b7d6-b28d76ef90f9 · outbound

This paper cites W., and Keutzer, K.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals W., and Keutzer, K

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.821257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.821257Z digest=sha256:e1b911d16e51f1ae4deb2b8a76d24c9cbc1300a03cdb1b29b339a9d26bb44487

Observation a187354c-1988-4510-bcea-69d363d219c8 · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.824457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.824457Z digest=sha256:75f95a765cf19c1dc87e7be9df08cea357442987513f8fc2a9742ef40bcdb04f

Observation b491a92e-362b-4c19-8fb1-d7eee7902bd1 · outbound

This paper cites APTQ : Attention-aware post-training mixed-precision quantization for large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals APTQ : Attention-aware post-training mixed-precision quantization for large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.874845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.828156Z digest=sha256:cd7d6d36a64edf1a7355046c70ff811f0205108e8d639b72904a1128c50a8d3e

Observation eb1cb692-bfda-4467-89e1-2f8988c49a39 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.831351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.831351Z digest=sha256:1362dd2fc498f3a3762ebd6382c7bd59fa380aeeb319621c8922a48baaf0d7a3

Observation 85af2621-dfc3-414b-8eb8-47ee1a297c0f · outbound

This paper cites Measuring massive multitask language understanding.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.835297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.835297Z digest=sha256:f9fc9029727f6575ed7dd07f6c53aceb1f8623f6989d2df40568fc095cfd2bc7

Observation 6aab0856-dddc-441b-bce0-76553e2988e1 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.838579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.838579Z digest=sha256:557b128e2c5c4d938becff68c36309313b2cd5ff20f0c581c2e339c21aba885b

Observation 94e41327-fd94-452c-9613-0033d004c734 · outbound

This paper cites SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.842294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.842294Z digest=sha256:608071fe38702d249082c6d0530b8b3be4c45f3e2b10af918250c0a844c3bd8b

Observation 180cdfcd-5a4f-449d-9476-ff9ccfa7946b · outbound

This paper cites Accurate post training quantization with small calibration sets.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Accurate post training quantization with small calibration sets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.856640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.845886Z digest=sha256:c000ca14be911539e5037c4a2088d8487ff6298436cc61661eaf3ee7d9a03419

Observation 8a50b5f7-d468-4e8b-bfbc-e2ed5fb7edd7 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.849655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.849655Z digest=sha256:c6ce3f7b7b5b2fe566e89d8b1c4ae41b00461fac716f18b0b73e70b50b456ee4

Observation 5194d093-8a59-44b6-8c34-a4f59fa04511 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.853192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.853192Z digest=sha256:68e02f1a6002f1a6c9e560da76f30186e1ff06946c9b06bd6e9c2bd29d7b31d7

Observation 01e71233-e6d8-40b0-8ba8-b8bf66f08fbf · outbound

This paper cites OWQ : Outlier-aware weight quantization for efficient fine-tuning and inference of large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals OWQ : Outlier-aware weight quantization for efficient fine-tuning and inference of large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.845115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.856813Z digest=sha256:073f7bf0c0d1695afb26e0958279dae92d408d1458fa9831e7243c43de237729

Observation 3670b04e-1dd8-4b6e-adfc-e19f7341568e · outbound

This paper cites SVDQuant : Absorbing outliers by low-rank components for 4-bit diffusion models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SVDQuant : Absorbing outliers by low-rank components for 4-bit diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.860219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.860219Z digest=sha256:eecda3633fa8e2893b2a8f2f8a20c908626ff15beeba59044e0c07e1769587f7

Observation b295b582-d3bf-47e4-b5c9-8ef67398a795 · outbound

This paper cites MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.863760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.863760Z digest=sha256:ec0254d7f49cfb572901186da9a80e5b12442991132b0e6167e2dbec24efcada

Observation 8518c039-3763-4a23-8a5d-9a0328f83243 · outbound

This paper cites Duquant: Distributing outliers via dual transformation makes stronger quantized llms.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Duquant: Distributing outliers via dual transformation makes stronger quantized llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.834633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.867491Z digest=sha256:bddb0d51021f8b1cc2072a3df86acd6e6b5f453fb14710aa7207e775d5d9215d

Observation f5882700-fdd5-418b-bdd4-ddb5ba719e1e · outbound

This paper cites AWQ : Activation-aware weight quantization for on-device llm compression and acceleration.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals AWQ : Activation-aware weight quantization for on-device llm compression and acceleration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.823746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.870930Z digest=sha256:173de4c0a23922d653f6052aafb1ade3714379a163e21a8309e37eb9c9f7dc20

Observation 672889be-de83-492a-8334-d73da03b3c3f · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.874322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.874322Z digest=sha256:56882665feee1973b931685cd2996d306b10bde5c1006412be06998e3cff2624

Observation 02159c26-b30b-4751-8b7a-edb233335d2d · outbound

This paper cites QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.878253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.878253Z digest=sha256:ad2027b234512e41f63637053fe3b601a8dfbd8658c9db800951075572e7a562

Observation a409a826-96ac-4f73-a143-d94ac889b68e · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.882488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.882488Z digest=sha256:3b69f5669b0eb1c4f74e580a58957d8742ade96d0a65e611f015bfa48444c85e

Observation a384ac31-e490-4143-920b-a8521ea8b37b · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.885979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.885979Z digest=sha256:462a328b70995fee0f9a578ec7a8ca9cf403539bbe302357830069c9416eb015

Observation ada04707-734b-4730-a56a-be7f57fb5790 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SpinQuant: LLM quantization with learned rotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.889642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.889642Z digest=sha256:2c6df5032eb74a888c54af34c02dddcea0f74105162bfba024ae14e1be331fcd

Observation 657afab0-bffc-4382-9b2d-d373b37bc0e7 · outbound

This paper cites Pointer Sentinel Mixture Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Pointer Sentinel Mixture Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.893145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.893145Z digest=sha256:a4edc6e4fbc914ee749a58dc38f218529c763856ba99d73802b8c7fce8e4fed6

Observation 7e525663-48d4-4e72-bfa3-8bf9528542f0 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models , 2024 a.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Llama 3.2: Revolutionizing edge AI and vision with open, customizable models , 2024 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.813238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.896642Z digest=sha256:c66f198628178ee9816934fcdec09ed7e8cb812908fdef98648685ef3f091885

Observation 71408979-c82a-4e96-84d7-4e1d34ae3197 · outbound

This paper cites Introducing Meta Llama 3: The most capable openly available LLM to date.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Introducing Meta Llama 3: The most capable openly available LLM to date

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.801562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.900006Z digest=sha256:66b6583945ae5d191fdeed2f36be4107e8bda3a72d5e763078d021772cdce6ce

Observation bbfbb140-674e-4a7a-bc1f-b9094136d1c0 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.903241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.903241Z digest=sha256:ecdf968ac6de3b816175d9ed52e89719eceb59a57f16ee0742e68e5612f29542

Observation 2abd8d63-e3fd-4201-bb72-1a8e6559d684 · outbound

This paper cites LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.906855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.906855Z digest=sha256:ea40493974af8cdf3893be3cdf8dc1a93e6e407b191caa4f07ef88c5ed13d73c

Observation 96cef960-686f-4279-9847-37ee47d223ca · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Pytorch: An imperative style, high-performance deep learning library

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.910471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.910471Z digest=sha256:8306f9d3788168a7e4ef4513b3a922b9abfb16989f62e9bb89a144d3d95c49c4

Observation 51a6c593-6f2c-4792-9f37-ba3063208c1f · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals L., Bhagavatula, C., and Choi, Y

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.777147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.913876Z digest=sha256:5db8e70732acab068425196ab173d2f9aa2e67e7c97b149ab96940d65d611ad7

Observation 2f89d0e0-1866-4a9c-be81-8e66a142ad62 · outbound

This paper cites ESPACE: Dimensionality Reduction of Activations for Model Compression.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ESPACE: Dimensionality Reduction of Activations for Model Compression

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:22:39.284687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.917322Z digest=sha256:60c2f349431f883a304349ba180c60b4d3853cb4f5d51a04316b8fffbed9d315

Observation cd559363-2a2b-4f94-977b-17763261db36 · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Social iqa: Commonsense reasoning about social interactions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.765854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.921067Z digest=sha256:0f38f2fa599e0001dd590925432c3a1d5aff6d988d28de23752b29aa2d9b3293

Observation 1ad94519-701c-4131-85b9-67b86ff644e7 · outbound

This paper cites Eigen attention: Attention in low-rank space for KV cache compression.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Eigen attention: Attention in low-rank space for KV cache compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.924318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.924318Z digest=sha256:dbbccef14bc3cd36efd07b0349b27426ae6d172048ae1d81c4ce199cae514feb

Observation 1c7b3e8d-46a8-40e0-ac33-5946b21554a6 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.927992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.927992Z digest=sha256:fd33986966a2e07ba3e14ddf8004bab5c73e504436b4822b423194aa5c4b7ac8

Observation bbdb286d-76dd-496f-968e-2e0daafb0ce1 · outbound

This paper cites Post training quantization of large language models with microscaling formats.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Post training quantization of large language models with microscaling formats

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.755113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.931852Z digest=sha256:3125a7f22b1b549c299cf1d180f3c5336df1b5e7d1d2c9574ec1f98d4a9c03e2

Observation c44c37e3-0c88-4b40-a9de-4b43cd10ed19 · outbound

This paper cites FlexGen : High-throughput generative inference of large language models with a single gpu.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals FlexGen : High-throughput generative inference of large language models with a single gpu

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.743866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.935507Z digest=sha256:4063a82c13c9fcf3f0e1af8c289ff9bd559122fd04e072153ba85a6584ea7851

Observation 0e18ae65-d51c-4214-8b61-b6d113c7b9a9 · outbound

This paper cites CUTLASS , January 2023.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals CUTLASS , January 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.732878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.938871Z digest=sha256:00617247a3d0d05cffc2c273648f148466c6e460f094d0fb2252adc8aceaa4e9

Observation ae6c5a15-afcb-4760-8200-43a06e826daf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.942577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.942577Z digest=sha256:52a930a898e96b08b4442768c1797b270d366864603da7b60dec1f0894b8bc83

Observation e6a3954c-345d-457a-a317-1c830cb69079 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.946066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.946066Z digest=sha256:a64a4dfb6b7bc92bfa0500b49e2ab447f20fd98e2aa22d7e956cd5dc5c63f1ba

Observation a24d496d-dfa2-4b16-9fb1-8f5f6267e687 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.949786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.949786Z digest=sha256:054f7dc0c389cd2920686cc00e195493e54d5d0c732a6d4fc9bebe48743b98c2

Observation f376dde5-0c87-4d47-9ce2-ec305c96d1ec · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.953613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.953613Z digest=sha256:89451bae0617c00fe1be4b323509caa81ea4e83ae09846cf5286811685b7f9d2

Observation b9d1323a-5e62-499c-b10c-edb95c4170a4 · outbound

This paper cites Training transformers with 4-bit integers.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Training transformers with 4-bit integers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.721622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.957266Z digest=sha256:74280ffa549bef48f90a7c71dc1b3d51738354ff35553677462c7299ac702cb6

Observation d1cd3c1c-d13b-4952-9c5e-cf547ed10e8d · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.960869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.960869Z digest=sha256:097322a0210a6665a4c1d1a19d263a21b03b8767af1d297454c4963a605f239d

Observation cefbaddf-3a47-45ac-a991-1b22181e5166 · outbound

This paper cites Qwen2.5 Technical Report.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Qwen2.5 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.964652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.964652Z digest=sha256:196ef56b40bc9e5db9fb08e2e189c0bbf69281f180f4f487df17ca41e12a540b

Observation 3ba31ebb-25e1-45fb-92bd-33a7d3d2f5d6 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.968304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.968304Z digest=sha256:cf13fe6fae6952acb259b1187ff39ed81b6ae9637648a15f21c307d45b62fa71

Observation e1b72ddd-8ca1-4837-8e6d-45c1e7ce89af · outbound

This paper cites ZeroQuant : Efficient and affordable post-training quantization for large-scale transformers.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ZeroQuant : Efficient and affordable post-training quantization for large-scale transformers

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.704547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.972203Z digest=sha256:88b0947bc96f00a21ec51cb8d707654e002941b0342a896a2b2319b9bf930cf0

Observation bfdc818f-8b94-49fe-811a-552d050a3e62 · outbound

This paper cites RPTQ: Reorder-based Post-training Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals RPTQ: Reorder-based Post-training Quantization for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.975957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.975957Z digest=sha256:0ee76f1aa26c20d074304f87cfeea7ef150e112cb48115a8c572d429073bb409

Observation 640b47c3-9b34-4981-8d9f-570908d32951 · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.979639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.979639Z digest=sha256:bfdfd3eed28a4d68958ea648e264d47bc3bfe21e2fc104faf0176599efd5d168

Observation 62b9c7e8-dff2-47be-8ce9-34a270e56684 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.983535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.983535Z digest=sha256:4c622396085eb6698f00d2d8103dc486b1612cb02fe58ec8f8a75b0d27066ed1

Observation 6768f08d-6eb8-4a9f-b10d-ea272f8c4cba · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.987961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.987961Z digest=sha256:6d08fc324ab1a1f836ccab7409674be9e2edbc0a4b3f8bbf310dc221eed1adc9

Observation e6ec41da-6b92-4b3d-97ed-afbeef829b46 · outbound

This paper cites ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.991673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.991673Z digest=sha256:d184a87bdb6ae7292729b11f5b1f52d230772dc431a329fb9ad71cecd3a9e3c4

Observation ba34eb41-461d-413b-98c4-71b9beb7e17e · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Atom: Low-bit quantization for efficient and accurate llm serving

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.693096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.995854Z digest=sha256:f1d575bc707945a694c5eaebe83fc46b75c64782077d1908d07fbefecdfe7192

Observation 82f4e82c-5f7c-4893-adf2-93f52d8046fc · outbound

This paper cites QMSum : A new benchmark for query-based multi-domain meeting summarization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QMSum : A new benchmark for query-based multi-domain meeting summarization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.681667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.999640Z digest=sha256:35166b436d7361c5c029647f4fd7f271ca91c39b356130939c122ac822537607

Observation 0b525f8a-5d78-4922-8207-1dc469f4a69c · outbound

This paper cites write newline.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals write newline

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:39.003914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:39.003914Z digest=sha256:be284696ba94bcf6078bae41db048917651c9480d287b9d5978679ee64628f0a

Pith citing papers

Observation 0989df39-3651-4f1f-b73d-8a0fd835d447 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.606027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.606027Z digest=sha256:198ec278e6d2d1c25a9972f58892abc7dbf1522e664d29ab93b0151fb2b40534

Observation 9354d518-b538-432f-b321-777feba70eff · inbound

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations cites this paper.

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T14:54:22.698198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:54:22.698198Z digest=sha256:749297c0f6438810a1c093e0e55b6ac14b8db45c85af4e05e73d7414a1c2e4fd

Observation db04ed1b-9a1f-47cd-889a-6f9c31fb8190 · inbound

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs cites this paper.

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T11:10:52.707779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:10:52.707779Z digest=sha256:e24ac68eca1b72fbbf0c236963c70a6fc98f2854fa676303c5522a8c66466436

Observation 203c50dc-9625-4f66-a193-588f56647cd9 · inbound

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales cites this paper.

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.544270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:44:42.989053Z digest=sha256:f59ce0673928b02ceacdb38cda7e796dcd8f58e534980259619febfa24d8b34e

Observation 2fa022f9-b499-44ef-8a7c-376882dc0c21 · inbound

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats cites this paper.

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.105639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T11:20:06.292977Z digest=sha256:7ffeb704f63133efe2287f618a867495a5d054d4110ffd072a58281110648f08

Observation 74e42f0f-80a9-4853-899d-dc176c6a364d · inbound

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats cites this paper.

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T10:56:54.155632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:56:54.155632Z digest=sha256:ea078d6d435009f8c2900a52017474b472aff0ba21a68612179e4e53d23740fc

Observation 76203019-c512-4b06-928c-48c9af433a7d · inbound

Multi-Bitwidth Quantization for LLMs Using Additive Codebooks cites this paper.

Multi-Bitwidth Quantization for LLMs Using Additive Codebooks ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:48:21.307051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T07:29:14.923431Z digest=sha256:f6a6fa5be5bde01086dcad3c7be4f24cc1f51edd0b47564329dd6f4aacd44860

Observation 0c189f19-d0d7-4bf7-96ca-0de477a6d0e3 · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:52.357889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:9be3bfd6cd0e5319f828b61e6dcbf5291677a8b85962dfa2bb9892b891748be0

Observation 040d99fe-e452-4a6e-8205-a5dc4e315599 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.892829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.892829Z digest=sha256:86e0eae3479da47b1aa392db361b844f9f7d3ece82a855e8250db6313785b39b