Pith. sign in

REVIEW 1 cited by

EPTQ: Enhanced Post-Training Quantization via Hessian-guided Network-wise Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.11531 v2 pith:VJCRWCLZ submitted 2023-09-20 cs.CV cs.LG

classification cs.CVcs.LG
keywords optimizationquantizationeptqnetwork-wiseboundmethodweightdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Quantization is a key method for deploying deep neural networks on edge devices with limited memory and computation resources. Recent improvements in Post-Training Quantization (PTQ) methods were achieved by an additional local optimization process for learning the weight quantization rounding policy. However, a gap exists when employing network-wise optimization with small representative datasets. In this paper, we propose a new method for enhanced PTQ (EPTQ) that employs a network-wise quantization optimization process, which benefits from considering cross-layer dependencies during optimization. EPTQ enables network-wise optimization with a small representative dataset using a novel sample-layer attention score based on a label-free Hessian matrix upper bound. The label-free approach makes our method suitable for the PTQ scheme. We give a theoretical analysis for the said bound and use it to construct a knowledge distillation loss that guides the optimization to focus on the more sensitive layers and samples. In addition, we leverage the Hessian upper bound to improve the weight quantization parameters selection by focusing on the more sensitive elements in the weight tensors. Empirically, by employing EPTQ we achieve state-of-the-art results on various models, tasks, and datasets, including ImageNet classification, COCO object detection, and Pascal-VOC for semantic segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A joint low-rank and mixed-precision quantization method that assigns rank and bit-width per transformer layer under a memory constraint, showing state-of-the-art compression accuracy.

Pith tools