Pith. sign in

Paper Citation Record · LEDGER

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2412.14590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14590 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T06:56:51.829741Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:38.647231Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:49:18.139563Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact14
  • verified fuzzy28
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37e62c2a-8bcb-451e-ae7e-285763dcdce8 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.068343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:d1e45d9f99e020be2441a2b28a24cd29f0ae490bc19a49d440469b920a59e25b

Observation 80233a95-9765-431e-8da9-c8aec65df4f4 · outbound

This paper cites Gulavani, Alexey Tumanov, and Ramachandran Ramjee.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Gulavani, Alexey Tumanov, and Ramachandran Ramjee

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.060729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:219355a74a871a3604e0d11d0cdc5f0fc4303c0fba2886ba188b4eab5488f973

Observation 3d05d206-a179-4663-894f-efe5f138528e · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.303236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:df4434cd6bd6f3785d4b5b72231963bdc7ffd370cfe4b8e5a40b0803e1179b02

Observation a79ce376-f19e-4d98-b145-ba71dbc37eb5 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.066006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:21e64dfac1823f70fd1aea9bb31bd2ff1515c80a423f6339f2c63a81bc64be79

Observation 9e2c2e61-c1ba-4ead-b6a2-fa5ad332ef80 · outbound

This paper cites Autogptq.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Autogptq

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.063348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:7a40e133c9231a51b05062e9e033e5f80a7d82f61b747addd25c0051d19fde65

Observation b435e830-b05f-4aed-a269-f03613b4b5c5 · outbound

This paper cites A systematic classification of knowledge, reasoning, and context within the ARC dataset.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design A systematic classification of knowledge, reasoning, and context within the ARC dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.076269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:e1650e3b28b310a6633e3f6cb1cd05fbc16078748a3a5981d83ff51f6f485cc8

Observation 328dbda5-aab3-4f19-9012-b69289a0ed10 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T06:57:40.294413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:517196ef28b3c62094f34561608eeea3b6955f4b62290663ba477e34067f52c7

Observation cf5f2a50-3df0-4687-92a0-c22a5f281592 · outbound

This paper cites Quip: 2-bit quantization of large language models with guarantees.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Quip: 2-bit quantization of large language models with guarantees

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.073634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a91d80472ff70dfac4a1cceb27fb2d821f247da9a3502709906da7fc394b64d9

Observation 44d0ce27-b02e-4ede-9552-4d15ae1f48d9 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.071072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:7a79915abd336525bb04b14f9053dd1caff6f459b3bb0c58e8dfd9d5af2918d1

Observation b567af91-f51d-4045-b786-727f64249083 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.298346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:1fa2f22db4a419c5afc7026dac9c3191fb6bb31cb37b69ee62459f999c05449c

Observation 36e87cfe-e0b8-4882-80f6-23765d942231 · outbound

This paper cites Spqr: A sparse-quantized representation for near-lossless LLM weight compression.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Spqr: A sparse-quantized representation for near-lossless LLM weight compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.133770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:8330c5b137cc6aabbc0762ac0227de3e234a87d69409aa598977f0bfdb413079

Observation 5dd118ca-c11b-4669-9c0a-5f7cfc8f4d10 · outbound

This paper cites Mahoney, and Kurt Keutzer.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, and Kurt Keutzer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.136046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:521ecab6f2bcc7e714ee35133d02ec8d1f0e04cf2b57ad8cd5a7a3f55cbfffa9

Observation 625e4c1c-2724-4088-9700-81c4b2dfb807 · outbound

This paper cites Optimal brain compression: A framework for accurate post-training quantization and pruning.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Optimal brain compression: A framework for accurate post-training quantization and pruning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.131468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:8ab7cd8c39bfc2cf2f19318f9b81901e31bede3d6895b1253f4ab48aac66730b

Observation 7d4685eb-7449-43f7-8289-f7c56f092f09 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.248123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:c0c97946565bd3d53bc6f7123af4c0809adaf12cd99bb45b753fbed96cec2227

Observation c7e285a5-b683-40a4-abc5-450fa0e3c48c · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design A framework for few-shot language model evaluation, 07 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.145836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:5d0b05e0f85b3bbff5647d835132b29a0f9a71b1c1e0cfe809736561bd4e9cb6

Observation 15924654-a6e9-4bb0-a154-96407715c746 · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design The Unreasonable Ineffectiveness of the Deeper Layers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.285477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:1aca3e35ff176837db7284306c268b78d0badddc605a9100977cc6804f36ee07

Observation 3c2d872c-6dc6-4506-bb6f-d8a2b6a34909 · outbound

This paper cites Qwen2.5: A party of foundation models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Qwen2.5: A party of foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.127122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:095b4083ee2c1d4f2197de3c8796cbff86b7a502ed329dc2f3a8e529209fbe55

Observation 2b793956-7161-48b8-8df3-de9b75fb91c2 · outbound

This paper cites Stork, and Gregory J.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Stork, and Gregory J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.129322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:9891d257e09bdbfe9c15dddaf35be869ec6be5a94d0c2a962e09f3688bc7a4c3

Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:75eb04107ca4e6d01d50e8b0ecf9f7da0cae75fd9b49376089379042a703e1d6

Observation dd189893-7f68-4131-ae81-4c7c4203e583 · outbound

This paper cites Mistral 7B.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mistral 7B

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.289405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:4f5a7b33b511ecf5b5496b5940255ce76a32f2ae3255f1e64b85281d30434096

Observation fccbe28f-94f4-442d-b8c0-fcc1435bc337 · outbound

This paper cites Mahoney, and Kurt Keutzer.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, and Kurt Keutzer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.140885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:f595b52ffbc0c9d54fd659d9a5f7d6869b0b34cf2d3890fc312ef9715bd1adee

Observation 7ebcf684-9784-4b84-b16c-b8a1dc19ec57 · outbound

This paper cites Scaling Laws for Precision.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Scaling Laws for Precision

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T06:57:40.270961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:34c8bf84874bc99bf4ab0790dd9bb99b57de210bc9c4ebe50e63367108533391

Observation 5d490239-93d0-46b6-87d2-8ccca3958243 · outbound

This paper cites Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.107546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:2c3df7b9f32162f02dfd3b60a3b6015eddebf0d83cc9f785cfa771c186632cbf

Observation 77a1fb6c-d6b1-455a-9e77-e80c3881cf1a · outbound

This paper cites Denker, and Sara A.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Denker, and Sara A

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.124803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:81b7769259eaebadf2f15559f8f0c0eea7c27f1898d02f3a615173b19614d49f

Observation 46af8c01-ddc2-46c9-8737-68ebd542122a · outbound

This paper cites OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.117685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:1fab7b5aa7a808682e9c1a160fbde0c7b645e9323df142484e5425e3950c5aa1

Observation 96ff1376-e8d0-45d4-bd22-8ccefc5e0f74 · outbound

This paper cites AWQ: activation-aware weight quantization for on-device LLM compression and acceleration.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design AWQ: activation-aware weight quantization for on-device LLM compression and acceleration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.119902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:4322d7ac2f581065790ec9c8c8405a57c57e9be210d803a54a6504e3efe83052

Observation 25cf0132-d301-4547-b25e-42fbae3358b2 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.275531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:adb37bf572159185f2eca5dacaf539a7143c1f38da55bd5e0da5dc54354e7a16

Observation 5cc3f7dc-0408-44ef-8b6a-a465057f4807 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design SpinQuant: LLM quantization with learned rotations

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.266681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:05d941838e155a589e332e4a61be205c32c830c6bbabb6aa8c2caa057bce3b22

Observation 082821a9-6969-4256-a07e-71703f945bdf · outbound

This paper cites Affinequant: Affine transformation quantization for large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Affinequant: Affine transformation quantization for large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.122100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:405dd20085036447c925e8bd0c8bd0495db428b4c918c944f7942d32f4119398

Observation c5fa5e41-febb-4a3d-acb1-bc99b492e866 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.238529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:8b22045c9c22e6b92b3097979d123a67e180d6f96f38527bccf3c59f90c70319

Observation 27706596-3357-4875-9fb2-efb2ad40f6f8 · outbound

This paper cites Pointer sentinel mixture models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Pointer sentinel mixture models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.109864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:f555093ffd3617cf570e332a97102f4c145ded3ad0b63c7c20432b74db451ce4

Observation 5925bd4b-8abd-4459-8de0-e2626acedd48 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.138711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:ec5bfb0dc7d28905c8601e756454616c9c011f91d2d589ce304bf211a94035c7

Observation 7ee67852-e4db-4291-b213-f86024dfce4a · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.112699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:511b07d937cfa51ad782a5d660ac0122374502379050a23f801ad76b4826ca52

Observation 24699ae9-ba64-4717-8032-0503701a09f8 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.115047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:7fcc97f5530df40dbeb75ceda4faf034ed5dffea15983b7d3621eec537422daa

Observation 70aa6d0d-76cc-418b-9d3e-b091e345afaf · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.105230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:5aad106ba46f4eb733b7d8b8dddd20ddb923528c5c653b2fa3f96266bdecbc8c

Observation df2d04e5-ba55-4ded-9db7-d4e2e52ad99f · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.233060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:fc778d998264b137fa17c927bb87c66c1f822bed4ba67ffd29719ce54d294aea

Observation 150ed859-5ba6-42ea-b95d-902aab744314 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Omniquant: Omnidirectionally calibrated quantization for large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.100063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:03eaff93436e5c223519302263d8f1b511103111aaa83a4adc30aec4bd60ea44

Observation 9a916006-04de-4242-a580-cce8adaaf96f · outbound

This paper cites Musr: Testing the limits of chain-of-thought with multistep soft reasoning.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Musr: Testing the limits of chain-of-thought with multistep soft reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.097254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:7acb9d5edfb0a667c94ac2dd38868c8fb76302d6cc2ed92a231f4a1c120ef1df

Observation 9ab5c030-5907-41f6-9520-ca5bbb367aad · outbound

This paper cites Le, Ed H.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Le, Ed H

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.102699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:1d097279a32b4c65fb090044376346c2db3d2497a1a5bbd893e622db959e0ed8

Observation 1af3be82-8754-4986-9e91-b642720089cf · outbound

This paper cites Tensorrt-llm.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Tensorrt-llm

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.091846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:07d9df83e466f0793b06b2410e4e1d012831bb5cd932e451ec1a44a7ea5a49d9

Observation 14bf5fc2-f624-4d6f-8bad-89df8cbd4141 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.280362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:fc4159f5b80040c3d0ca92d4f5350008621b7db53039d6659abd7dfee50dde8e

Observation eed79741-d8c4-4379-8170-d62df4d12977 · outbound

This paper cites ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.258320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:543788a57fb4d14a9a00b2b5264686ae3ddebda3f18499360c3d523939b4dace

Observation 76c4fb00-95ce-4142-b113-fc6415d154b3 · outbound

This paper cites Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.094814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:84421fe813f16febbc47c88ff67e876c657f6e6099cfff1ed2587d7c21cc2703

Observation a272b44b-2ee9-4610-80eb-f1e7cca57fee · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.143164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:5b9a108fbdc04d4d131fe6e45ae064d52b2de240ec3683432572bae2d47017bd

Observation 852e474e-3a66-475d-af0f-0d40db242d5b · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.087377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:bc09907dafee6a5261a81d7ae2cf90fdc5a480607b6436310234a1bb38edc044

Observation b98ecc84-49b8-424c-806e-381e79472077 · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.084338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:f66ac7d59e72e44bd2f91f7f81513e20bf490600003d784309ca9c08252be236

Observation 06da1fc8-e47f-4710-8f4f-da7a2c3af0d9 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Orca: A distributed serving system for transformer-based generative models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.089673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:e71863d4e711c0ca5141eb2e4626f1f3e25044e34feb05f08960deae9dd0dcdc

Observation f4de0dba-4110-4ca8-9826-474856e95ec3 · outbound

This paper cites RPTQ: Reorder-based Post-training Quantization for Large Language Models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design RPTQ: Reorder-based Post-training Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.243433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:89ea7b7f4bb9c6bec1801ca4824911d875008848169970441a69e2bee3f75089

Observation 1c153e90-38a1-464b-9575-64af6d836d91 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.079285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:56b184356f84c8e99c46e35221b035b5c0fa077fdf723ea2034608bb98dc98cd

Observation 288a7325-3656-4722-bda6-b4afa6948d63 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate LLM serving.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Atom: Low-bit quantization for efficient and accurate LLM serving

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.081983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:9451671fc54dd81d981418d449fe59c18e498a6dc5c32ef75fee8574d876e709

Observation 3c7fae06-d6b4-4e1a-8f53-58f68161af7f · outbound

This paper cites BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.253161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:bdeb65193ed98970929bd377d7920285dacd8e8e4f31c779a0b6df3b6f59f4f0

Pith citing papers

Observation 2485f3ff-a7bd-485d-a157-0c9e71abb003 · inbound

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method cites this paper.

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-06T14:45:38.647231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:38.647231Z digest=sha256:5ac52f41ee7f8ea749d759710fc443422ddd2584c2505aaf06bbf882f07b43be

Observation f7421b5f-a3f4-4fec-90df-40557947d13d · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:23:03.651022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:6b9a2c5116ec40017996615c75e16094802a42c1b31290c1c92c3ecc9203a7eb

Observation 357fa1c1-a011-493d-b621-bca2b712eb82 · inbound

Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment cites this paper.

Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:49:18.141527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T20:58:33.720040Z digest=sha256:82f44977f723a0bb5d5ccd7dbe14d05ccce90e8416424620dba3256da94ecf42

Observation 58e27d8e-bafb-43ad-9522-5e1bd0616350 · inbound

Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models cites this paper.

Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T06:20:07.112455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:20:07.112455Z digest=sha256:af3999c9cc619985015b2fba1136116eeec409ea082cc740c4bfb926ddb33c03