Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:30.222315Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.01144.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:30.222315Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T19:31:28.053090Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
36 of 36 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 59c233a4-120d-470f-b98c-b007632f82a6 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5730738b-99fe-49ed-a07a-4295aaab48a5 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de82ca2d-16c6-454f-93cd-0ebcd8d9434d · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6845930e-22d5-46b2-b580-465d4fbe2a77 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c5d3b4-f15c-4449-8cb2-7a52421e01cb · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Extreme Compression of Large Language Models via Additive Quantization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44839155-e3cc-4ca0-9edd-e8f9df445a1a · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd3db42-e086-469b-afd7-86e140747b2e · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Measuring Massive Multitask Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12508ec8-b29a-4e3d-b6eb-f5b092297b26 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d957495-3395-43b2-b49d-98a464d2b749 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 845afe81-1327-4b44-ae0e-bd8772d76f91 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference FPTQ: Fine-grained Post-Training Quantization for Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab32f4d4-f661-4a35-966a-79d4cfc4a8fe · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM-FP4: 4-Bit Floating-Point Quantized Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd60780-c798-4c51-a749-3164d5e1d12e · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 620618b6-9e81-41a0-acc5-3d44f4f5853e · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Pointer Sentinel Mixture Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a922f6-ae9a-4391-9fc2-12d018e30109 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Microscaling Data Formats for Deep Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e330fa-b746-4e55-8786-03d9f1dd90ed · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa2e2a66-609a-43b6-91b9-f657944f403b · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Release Strategies and the Social Impacts of Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b41623d-4886-43d5-a8e0-207e4be94fc1 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8116dac2-c1ee-4493-a86f-e24760ad3bdb · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9ee65de-9815-4904-a74b-f6315e105749 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b644334-71ea-409b-bab9-a42091015223 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae98df9-0dae-4ba3-826d-c91b58bb71fe · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 443379e9-08ad-40c8-a500-52f77b1244a7 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08012266-ae3c-495e-b919-96016dba2b7f · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a341f18-8111-494f-8508-5b9ca07d02e4 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BlockDialect quantizes matrices and vectors along their respective multiplication dimensions
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc91eacd-7966-4cd4-9cd6-af15ba681bde · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference These approaches often dequantize data to FP16 before performing multiplications, which limits computational efficiency
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 291a412a-29bd-46d2-bdf7-84eb21efe9df · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fe82128d-1d2a-4c50-8bde-56c642954bb2 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Infinitesimal generators for a class of polynomial processes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0666a55e-aba5-4c73-a525-14dc7910b484 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference However, 2D block quantization generally results in higher perplexity
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56011b9a-eab4-4f73-8e41-f7b303c754cc · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f44717-6eb3-46df-80a9-6fa3c8c7ed8b · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4009a59-93ff-46a6-a113-6d43629015c9 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e256fefb-690f-4ec1-9063-af310016def5 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba4d08ae-e315-4153-882f-88fc32a473a4 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c86a6e86-becd-4432-97fa-457a4d37d62f · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e991ba-1024-466c-94aa-34eb68889dd3 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdbfc1e9-7671-4337-a652-aaff7238b9d6 · outbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7fbe63-008d-49dd-99cb-af5d1d277a18 · inbound
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.