Pith. sign in

REVIEW 13 cited by

Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.09602 v1 pith:4CAQC7FG submitted 2020-04-20 cs.LG stat.ML

classification cs.LGstat.ML
keywords quantizationintegerdeepincludinginferencemodelsnetworksneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Quantization techniques can reduce the size of Deep Neural Networks and improve inference latency and throughput by taking advantage of high throughput integer instructions. In this paper we review the mathematical aspects of quantization parameters and evaluate their choices on a wide range of neural network models for different application domains, including vision, speech, and language. We focus on quantization techniques that are amenable to acceleration by processors with high-throughput integer math pipelines. We also present a workflow for 8-bit quantization that is able to maintain accuracy within 1% of the floating-point baseline on all networks studied, including models that are more difficult to quantize, such as MobileNets and BERT-large.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 218 citations worldwide. Full citation record

  1. Jack of All Scales: A Versatile FPGA Tensor Block for MXFP Precisions

    cs.AR 2026-07 conditional novelty 7.0 of 10

    Targeted DSP tensor-mode changes enable native MXFP4/MXFP6/E4M3 support on Agilex-5-like FPGAs, with a 36% block-area cost and 4.2x average systolic-array throughput gain over baseline mapping strategies.

  2. INT8 Quantization Makes ARM Edge Inference Dispatch-Invariant

    cs.ET 2026-07 conditional novelty 6.0 of 10

    INT8 QDQ CNNs are byte-exact across ARM Cortex-A53/A72/A76 under XNNPACK even with different SIMD kernels, unlike FP32 or x86 INT8.

  3. Float8@2bits: Entropy Coding Enables Data-Free Model Compression

    cs.LG 2026-01 conditional novelty 6.0 of 10

    EntQuant stores LLM weights at ~2 bits per parameter by entropy-coding Float8 weights, matching data-dependent compression quality without needing calibration data.

  4. FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models

    cs.LG 2025-08 conditional novelty 6.0 of 10

    FlashSVD fuses low-rank SVD projections into attention and feed-forward GPU kernels so SVD-compressed transformers avoid materializing dense activations, cutting activation memory at a latency cost.

  5. DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A post-training quantization method that combines learned channel scaling and power-of-two scaling keeps diffusion image quality high at 4-bit weight, 6-bit activation precision.

  6. Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition

    cs.LG 2025-06 conditional novelty 6.0 of 10

    ODLRI initializes the low-rank component using activation-outlier channels, improving low-bit compression of large language models over the CALDERA baseline.

  7. Adaptive Semantic Token Communication for Transformer-based Edge Inference

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A single adaptive deep joint source-channel coding model with budget-conditioned token selection and Lyapunov-based resource allocation achieves better accuracy-compression trade-offs than static DJSCC and digital bas...

  8. Approximate reservoir computing with a semiconductor laser for reducing energy consumption

    physics.optics 2026-07 conditional novelty 5.0 of 10

    Quantizing laser-based reservoir computing to 5 bits and lowering sampling frequency and drive current cuts estimated energy per sample by 91% without losing prediction accuracy.

  9. Real-Time Analysis of Unstructured Data with Machine Learning on Heterogeneous Architectures

    physics.data-an 2025-08 conditional novelty 5.0 of 10

    A graph neural network (ETX4VELO) reconstructs LHCb VELO tracks with performance comparable to the production 'search by triplet' algorithm while running end to end in the GPU-based first-level trigger, with additiona...

  10. Compress Any Segment Anything Model (SAM)

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Birkhoff compresses 18 SAM variants to about one-fifth their size with less than 1% accuracy loss, data-free, using a trajectory-based codebook and a fused GPU operator.

  11. Efficient EEG Seizure Detection Using INT8 Quantization, Channel Pruning, and Spiking Neural Networks

    eess.SP 2026-07 conditional novelty 4.0 of 10

    On a shared 1D-CNN baseline for CHB-MIT seizure detection, INT8 quantization cut model size from 1.63 to 0.44 MB and latency by 2.8x with preserved AUC, while SNN conversion was 288x slower on CPU.

  12. Enhancing Automatic PT Tagging for MEDLINE Citations Using Transformer-Based Models

    cs.DL 2025-06 reject novelty 4.0 of 10

    Transformer-based models can predict many MeSH Publication Types with high accuracy, but the paper does not fairly compare them to the existing MTI baseline, leaving the central improvement claim unsupported.

  13. Power-of-Two (PoT) Weights in Large Language Models (LLMs)

    eess.SP 2025-05 conditional novelty 4.0 of 10

    Power-of-two weight quantization applied post-training to a 124M GPT-2 model degrades cross-entropy from 3.17 to about 4.1-4.5 at 4-6 bits, while promising memory and bit-shift savings.

Pith tools