Pith. sign in

REVIEW 8 cited by

Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.05662 v4 pith:K62EWTFO submitted 2019-03-13 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords gradienttraininglossactivationchaincoarsepopulationrule
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training activation quantized neural networks involves minimizing a piecewise constant function whose gradient vanishes almost everywhere, which is undesirable for the standard back-propagation or chain rule. An empirical way around this issue is to use a straight-through estimator (STE) (Bengio et al., 2013) in the backward pass only, so that the "gradient" through the modified chain rule becomes non-trivial. Since this unusual "gradient" is certainly not the gradient of loss function, the following question arises: why searching in its negative direction minimizes the training loss? In this paper, we provide the theoretical justification of the concept of STE by answering this question. We consider the problem of learning a two-linear-layer network with binarized ReLU activation and Gaussian input data. We shall refer to the unusual "gradient" given by the STE-modifed chain rule as coarse gradient. The choice of STE is not unique. We prove that if the STE is properly chosen, the expected coarse gradient correlates positively with the population gradient (not available for the training), and its negation is a descent direction for minimizing the population loss. We further show the associated coarse gradient descent algorithm converges to a critical point of the population loss minimization problem. Moreover, we show that a poor choice of STE leads to instability of the training algorithm near certain local minima, which is verified with CIFAR-10 experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MusicMark: A Robust Generative Watermarking Framework for Music Generation

    cs.SD 2026-07 conditional novelty 6.5 of 10

    Embedding watermark bits into diffusion semantic latents via a frozen-backbone adapter yields far more robust music provenance than post-hoc watermarking under codecs and cover-song attacks.

  2. DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and Optimization

    cs.AR 2026-08 conditional novelty 6.0 of 10

    A GPU-powered differentiable switching power engine uses bytecode automatic differentiation and simulation-based calibration to deliver fast power gradients for cell sizing and power virus search.

  3. Beyond Token-Level Cross-Entropy: Fr\'echet Distributional Post-Training for Autoregressive Image Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    FD-loss post-training with detached rollout replay and a probability-level straight-through estimator improves FID and FD_r6 across eight ImageNet configurations.

  4. Energy Efficiency Maximization for Hybrid RIS-Aided Communications via Deep Unfolding

    eess.SP 2026-07 conditional novelty 6.0 of 10

    Deep-unfolded alternating optimization for hybrid RIS mode selection and binary phases yields ~30% higher energy efficiency than plain projected gradient and most of the gain with only ~10% of elements active.

  5. Neural-Network-Assisted Binary Template Construction for Matrix-Based Pattern Matching in the STCF MDC

    hep-ex 2026-07 conditional novelty 6.0 of 10

    Offline multi-objective neural optimization yields compact binary trigger–recovery template pairs that retain ~98% of true MDC hits at 95% detector efficiency under multi-background conditions without changing the onl...

  6. Synesthesia of Machines (SoM)-Aided LiDAR Point Cloud Transmission for Collaborative Perception

    eess.SP 2025-09 conditional novelty 6.0 of 10

    LPC-FT transmits LiDAR point clouds as compressed learned features over digital channels and, in OpV2V experiments, reduces Chamfer Distance by 30% and raises PSNR by 1.9 dB versus the SEPT baseline.

  7. TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision

    cs.LG 2025-06 conditional novelty 6.0 of 10

    TruncQuant uses a floor-based quantizer with 2^n scaling instead of rounding with 2^n-1, so truncating high-precision weights via bit-shift exactly matches direct low-precision quantization, recovering accuracy lost b...

  8. RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    RoSTE couples quantization-aware supervised fine-tuning with per-layer Hadamard rotation selection, reducing quantization outliers and improving 4-bit quantized LLM accuracy over SFT-then-PTQ baselines.

Pith tools