Pith. sign in

Paper Citation Record · LEDGER

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference

As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.01144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01144 v5

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:30.222315Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T19:31:28.053090Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 59c233a4-120d-470f-b98c-b007632f82a6 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.059448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.059448Z digest=sha256:990c81d64bed3c44b1867fe39a82be4cc383b57992a1ffb5d456cf21474f0736

Observation 5730738b-99fe-49ed-a07a-4295aaab48a5 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.075869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.075869Z digest=sha256:741b36bfdeeff739302c2a89d7dcab6c848a00f87cd20c2182dd93239ed71cb7

Observation de82ca2d-16c6-454f-93cd-0ebcd8d9434d · outbound

This paper cites Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:44:30.653784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.091674Z digest=sha256:7d738dd150456e75c3a5f2714b95c715de9268cac90755fe107f8682d58ccb00

Observation 6845930e-22d5-46b2-b580-465d4fbe2a77 · outbound

This paper cites The Llama 3 Herd of Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.097173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.097173Z digest=sha256:4730e020104acbb85d8ff73bbd629fae558faed26ad9b698332dd9260ddd4369

Observation 66c5d3b4-f15c-4449-8cb2-7a52421e01cb · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Extreme Compression of Large Language Models via Additive Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.102307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.102307Z digest=sha256:212a9f8bb1195989041cf13b4461f86bb04685adc069b6704565aba1aa8c5459

Observation 44839155-e3cc-4ca0-9edd-e8f9df445a1a · outbound

This paper cites BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.107005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.107005Z digest=sha256:d310cadef9b09049db984191cf5137f8144f24324d31c818f9f90c9f97ca76ad

Observation abd3db42-e086-469b-afd7-86e140747b2e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.115643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.115643Z digest=sha256:fd89e36b17cb8e57902e26624d4ab24478756bb1476627855fa527681abacd0c

Observation 12508ec8-b29a-4e3d-b6eb-f5b092297b26 · outbound

This paper cites Mistral 7B.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.124358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.124358Z digest=sha256:91d8f9919c5951755b8e79a0231aaea91b59558e421f388afaeb405c6ec52ab4

Observation 7d957495-3395-43b2-b49d-98a464d2b749 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.128860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.128860Z digest=sha256:bab0b2178dfd857fd59e310541f455faabbaaf3c8cf47887d5ffaee6e869d221

Observation 845afe81-1327-4b44-ae0e-bd8772d76f91 · outbound

This paper cites FPTQ: Fine-grained Post-Training Quantization for Large Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference FPTQ: Fine-grained Post-Training Quantization for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.133415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.133415Z digest=sha256:7c8fd7705e39b8ba99e14b95ec3c605f5ba108ae0c9ee5832904ffcb1946c5f4

Observation ab32f4d4-f661-4a35-966a-79d4cfc4a8fe · outbound

This paper cites LLM-FP4: 4-Bit Floating-Point Quantized Transformers.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM-FP4: 4-Bit Floating-Point Quantized Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.137644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.137644Z digest=sha256:bc8945f4802fcd7f23be90ebe700b60f1a8ab411ac4e274d7830fe7891febae0

Observation 6fd60780-c798-4c51-a749-3164d5e1d12e · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.142155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.142155Z digest=sha256:321c8689d8de5047fa3fcd601fad5e1a33d05ac7659d48f1224b541ef3d5de92

Observation 620618b6-9e81-41a0-acc5-3d44f4f5853e · outbound

This paper cites Pointer Sentinel Mixture Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Pointer Sentinel Mixture Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.147521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.147521Z digest=sha256:56a69227a3eaf27d73472325404bac97fab0e159d0d17ec5efb9dcd988b1ddc5

Observation 30a922f6-ae9a-4391-9fc2-12d018e30109 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Microscaling Data Formats for Deep Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.157115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.157115Z digest=sha256:ca3955409b11fcdac2779cf821695f445a1e2ff61ea222a1a1edd7b2caf49ed9

Observation 45e330fa-b746-4e55-8786-03d9f1dd90ed · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.161918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.161918Z digest=sha256:38f747c5c3a8315a36e1e9ab4f49fadce852db649ad9e5bcb0b40d6482b8be71

Observation fa2e2a66-609a-43b6-91b9-f657944f403b · outbound

This paper cites Release Strategies and the Social Impacts of Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Release Strategies and the Social Impacts of Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.167061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.167061Z digest=sha256:9ae48317035115ad20820b4d43f37d8e6139a327fc7d46293d7afee19ee368c0

Observation 6b41623d-4886-43d5-a8e0-207e4be94fc1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.171823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.171823Z digest=sha256:016d122ea2906be94e5e3d17fcf2f005ec38b656bde640833c2e53674c73275b

Observation 8116dac2-c1ee-4493-a86f-e24760ad3bdb · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.175978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.175978Z digest=sha256:08d920f06db664982dd00b91a7b1c137fe68b50f8620405fe9577194cc87bd21

Observation c9ee65de-9815-4904-a74b-f6315e105749 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.180277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.180277Z digest=sha256:a32e6ce07c21c8eee4999bc4282093d51a8746d5329627f0777014399ce819ff

Observation 8b644334-71ea-409b-bab9-a42091015223 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.189163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.189163Z digest=sha256:4b301efdb62641bc08817da2658233891a5cc59d4e3e1a413e13c0f2d55212ce

Observation bae98df9-0dae-4ba3-826d-c91b58bb71fe · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.193758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.193758Z digest=sha256:a5d74bdc276148920390ccaacaa220125e94f8051876424ba03b5499c77fc921

Observation 443379e9-08ad-40c8-a500-52f77b1244a7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.198608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.198608Z digest=sha256:ca9e3e50b17a9e59352ff31d65d7a51441419edee51aa206f1d1a9d9b88179f2

Observation 08012266-ae3c-495e-b919-96016dba2b7f · outbound

This paper cites Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.792866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.203582Z digest=sha256:54f8727cc4b0cbc28e76aba101f18ea53e7ac50f42ec46e068c83ad66bd2d23b

Observation 1a341f18-8111-494f-8508-5b9ca07d02e4 · outbound

This paper cites BlockDialect quantizes matrices and vectors along their respective multiplication dimensions.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BlockDialect quantizes matrices and vectors along their respective multiplication dimensions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.779589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.207612Z digest=sha256:14ab518948651d2dde1a9c0a314d45c1b69d19dfde9c7ed532b2e6c09ed89f90

Observation bc91eacd-7966-4cd4-9cd6-af15ba681bde · outbound

This paper cites These approaches often dequantize data to FP16 before performing multiplications, which limits computational efficiency.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference These approaches often dequantize data to FP16 before performing multiplications, which limits computational efficiency

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.766708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.211562Z digest=sha256:8c2d3b8f21dfc68ee804d9269b50391002af85d7934ca349ae8a0816589ec7b7

Observation 291a412a-29bd-46d2-bdf7-84eb21efe9df · outbound

This paper cites an unresolved cited work.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Unresolved cited work

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:44:30.752293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.215113Z digest=sha256:7610443bbb167c8af6b34d994706ede0146104ca6d75f5126c488652bcbc20e9

Observation fe82128d-1d2a-4c50-8bde-56c642954bb2 · outbound

This paper cites Infinitesimal generators for a class of polynomial processes.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Infinitesimal generators for a class of polynomial processes

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T22:44:30.261174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.218614Z digest=sha256:98f006321658c8d9dbee3389a5ef515dccb5d8af127d2b5eae766cf8ea04a646

Observation 0666a55e-aba5-4c73-a525-14dc7910b484 · outbound

This paper cites However, 2D block quantization generally results in higher perplexity.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference However, 2D block quantization generally results in higher perplexity

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.739617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.222315Z digest=sha256:b4bbcfd711e64c705b1d09475f03aa22504bd81c1184dd54c6e6aa5f6d4d6490

Observation 56011b9a-eab4-4f73-8e41-f7b303c754cc · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.152331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.152331Z digest=sha256:f95719711b9d34d68eaf2782c247c589a9ed1c62b2dc442810aca12a6a8da8bf

Observation 36f44717-6eb3-46df-80a9-6fa3c8c7ed8b · outbound

This paper cites ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.184926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.184926Z digest=sha256:b5f74a6c80abb9e2dff7bab34ac823f4d64941935e21430bdaa5e16c19dd07e8

Observation b4009a59-93ff-46a6-a113-6d43629015c9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.081186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.081186Z digest=sha256:c6c8d76bd8fa70769bfe3fc3b164e432c9b7100db93b07192cf7556abce5ab48

Observation e256fefb-690f-4ec1-9063-af310016def5 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.120356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.120356Z digest=sha256:7003a342ce3725b173e7915c61ecef400676cfd5aabd70f80fcdb5f33f8686bf

Observation ba4d08ae-e315-4153-882f-88fc32a473a4 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.086096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.086096Z digest=sha256:eede7d5f56283a666e9669390406b05ed96413662a7fd5967e944dfc049a815f

Observation c86a6e86-becd-4432-97fa-457a4d37d62f · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.071130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.071130Z digest=sha256:e8b4f5fa88246068e5745a03847b179b21fa55e57b8503290e8cbc3f8718bfd1

Observation 46e991ba-1024-466c-94aa-34eb68889dd3 · outbound

This paper cites QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.065512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.065512Z digest=sha256:7eba3a9a5fb30ab136b1fe362950113e35275ce4b634b377de9cda99a9dc1cda

Observation cdbfc1e9-7671-4337-a652-aaff7238b9d6 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.111333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.111333Z digest=sha256:a94210019eecc7f7ca7ac245cfc4c3934b9d1897b3211da991d765fd4bd6a2e3

Pith citing papers

Observation fe7fbe63-008d-49dd-99cb-af5d1d277a18 · inbound

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference cites this paper.

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:34.309836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T19:31:28.053090Z digest=sha256:a281d52bd3a3dec485382c40e0209b9141ed2670ac55c540d369e663a33a8ec6