Pith. sign in

REVIEW 5 major objections 7 minor 78 references

Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations

T0 review · 5 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Low-bit convolutional networks can be trained to beat their quantized baselines.

desk verdict Useful quantization training tricks with a real confound: SP gains may be training-budget artifacts, but KD and progressive quantization hold up. read the letter →

arxiv 1908.04680 v3 pith:SXMJKKYO submitted 2019-08-10 cs.CV

classification cs.CV
keywords low-bitwidthneuralnetworksmodelquantizationprogressivestochasticprecisionknowledgedistillationstraight-throughestimatorimageclassificationconvolutional
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the accuracy drop of convolutional networks with very low-bit weights and activations is mainly a training problem, not a capacity problem, and that it can be mitigated by changing how the network is optimized. It proposes three training strategies: progressive quantization, which solves easier subproblems first by quantizing weights before activations and by gradually lowering bit-width; stochastic precision, which randomly quantizes only part of the network each iteration while annealing the quantized fraction to the full model; and joint knowledge distillation, in which a full-precision teacher and a low-precision student are trained together and adapt to each other. The central evidence is that these strategies, alone or combined, lift 2-bit ResNet-50 on ImageNet from 70.19% to 73.25% top-1 accuracy over the DoReFa-Net baseline. If true, this matters because low-bit networks offer large memory and compute savings, and the proposed training techniques are orthogonal to the specific quantizer, so they could be combined with other quantization schemes.

What carries the argument

The machinery is the training schedule itself, built on three mechanisms. First, a fixed-point quantizer of the DoReFa form, $Q(z) = \frac{1}{2^k-1}\mathrm{round}((2^k-1)z_r)$, is combined with a straight-through gradient approximation, $\partial z_q/\partial z_r \approx 1$, so gradient-based updates can flow through a non-differentiable quantizer. Second, a binary mask over network fragments (layers or residual blocks, and optionally weights versus activations) selects which parts are quantized at each iteration, with the stochastic ratio $\delta$ starting at 0.5 and linearly decaying to zero; this provides the progressive-relaxation effect in one stage. Third, a two-headed loss couples the low-precision student to a jointly updated full-precision teacher through KL divergence on posterior probabilities, weighted by $\beta = 0.5$, and attention transfer on feature maps, weighted by $\gamma = 50$. These mechanisms carry the argument by making the discrete optimization tractable and by continuously guiding the student toward the teacher during training.

What would settle it

Re-run the 2-bit ResNet-50 ImageNet experiment with the paper's exact hyperparameters but disable stochastic precision while keeping joint distillation: the paper reports 71.96% top-1 without SP and 73.25% with it, so a replication showing no statistically significant gap would falsify the claim that stochastic precision adds a real benefit; similarly, disabling the distillation losses should reproduce the 70.19% baseline if the teacher guidance is genuinely responsible for the gain.

Watch

Extended reading notes

Core claim

The paper's central claim is that the optimization difficulty of quantized neural networks, not the capacity of the low-precision representation, is the main obstacle to accurate low-bit models, and that this difficulty can be attacked with three training techniques. Progressive quantization first solves the easier problem of quantized weights with full-precision activations, then adds activation quantization, and also anneals the bit-width from 32 bits down to the target precision. Stochastic precision randomly quantizes a fraction of layers, blocks, weights, or activations per iteration while the stochastic ratio is linearly decayed to zero, easing gradient flow in a single training stage. Joint knowledge distillation trains a full-precision teacher and a low-precision student together, using both posterior KL divergence and attention-transfer losses, and reports that the student surpasses the baseline and sometimes the teacher improves as well. The paper reports that all low-precision models produced by these methods surpass their corresponding baselines, with the largest combined gain reaching 73.25% top-1 accuracy for 2-bit ResNet-50 on ImageNet.

Load-bearing premise

The load-bearing premise is that the straight-through estimator supplies genuinely useful gradient directions at very low precision, and that the hand-set schedule of stochastic ratios and distillation weights keeps working for other architectures, datasets, and quantization schemes.

Editorial extensions

If this is right

  • 2-bit ResNet-50 on ImageNet can reach 73.25% top-1 accuracy with stochastic precision plus joint distillation, compared with 70.19% for the DoReFa-Net baseline, closing most of the gap to the 75.64% full-precision model.
  • The three training strategies are complementary: two-stage optimization, progressive precision, stochastic precision, and knowledge distillation can be combined, and each combination in the paper improves over the corresponding baseline.
  • Joint distillation helps more when the low-precision student is harder to train, since learning from scratch shows larger relative gains than fine-tuning, and quantized ResNet benefits more than quantized PreResNet.
  • The training techniques are orthogonal to the quantizer: stochastic precision improves DoReFa-Net, LQ-Net, BiReal-Net, and GroupNet baselines, suggesting the methods could be transferred to other quantization designs.
  • The full-precision teacher can also be slightly improved by joint training with the student, indicating that the mutual adaptation acts as a regularizer rather than a one-way transfer.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension, not pursued in the paper, would be to replace the hand-set stochastic ratio decay and distillation weights with learned or per-layer schedules, since the paper gives no sensitivity analysis for these hyperparameters and the gains might be further improved or made more robust.
  • Beyond image classification, one could test the same three training strategies on low-bit object detection or semantic segmentation, where the optimization difficulty is typically greater and the reported benefits might be even more pronounced.
  • The observation that the jointly updated teacher improves suggests that mutual distillation could serve as a general regularizer for full-precision training, independent of quantization.
  • Because the methods are quantizer-agnostic, a testable extension is to pair them with modern learned-step-size quantizers or with mixed-precision allocation, potentially recovering additional accuracy at the same average bit-width.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper proposes three training strategies for low-bitwidth CNNs: (i) progressive quantization, with a two-stage scheme (quantize weights first, then activations) and a progressive precision scheme (32→8→4→2 bits during training); (ii) stochastic precision (SP), which randomly quantizes fragments of the network while keeping the rest full-precision, with the stochastic ratio decaying to zero; and (iii) joint knowledge distillation (KD), in which a full-precision teacher and a low-precision student are trained together using both KL-divergence and attention-transfer losses. Experiments on CIFAR-100 and ImageNet with DoReFa-Net, LQ-Net, BiReal-Net, and GroupNet report consistent accuracy improvements over the corresponding baselines, culminating in combined SP+KD results such as 73.25% top-1 for 2-bit ResNet-50 versus a 70.19% baseline.

Significance. If the reported gains hold under controlled comparisons, the paper would make a useful practical contribution: it offers simple, quantizer-agnostic training recipes that improve low-bitword accuracy across several architectures and quantization schemes, and the combination of stochastic precision with joint distillation appears complementary rather than redundant. The paper also ships a fairly large set of experiments with repeated runs and standard deviations for most of the headline tables, which is a strength. However, the central quantitative claim is weakened by a training-budget confound in the SP and SP+KD experiments, by apparent typographical inconsistencies in Table 1, and by the absence of sensitivity analysis for the key hyperparameters (δ, β, γ, and the precision schedule). The contribution is incremental over the authors' CVPR 2018 conference paper, but the new one-stage SP formulation and the joint-teacher distillation study constitute a reasonable extension for a journal version.

major comments (5)
  1. [Sec. 4.2 and Table 2] The SP experiments use a 40-epoch schedule (learning-rate decays at epochs 25 and 35), while the baselines are trained for 30 epochs with decays at epochs 15 and 25, as stated in Sec. 4. The baseline numbers in Table 2 and Table 9 are identical to those in Table 4, which uses the default 30-epoch schedule. No baseline re-run under the 40-epoch schedule is reported, so the gains attributed to SP (e.g., 70.19→72.23 in Table 2) and to SP+KD (70.19→73.25 in Table 9) are confounded by the additional 10 epochs of training and the later learning-rate decays. The authors should either re-run the baselines under the SP schedule or provide evidence, such as a 40-epoch baseline curve, that the extra epochs alone do not explain the improvement. Without this, the central claim that 'all our low-precision models surpass the corresponding baselines' is not established on a like-for-like training budget.
  2. [Table 1] The top-5 accuracy entries for the 4W,4A ResNet-50* row are internally inconsistent: the Baseline is listed as 75.70, which is below the Top-1 value of 75.11, and is much lower than the Baseline+TS value of 91.93. Similarly, the 2W,2A ResNet-50* Baseline top-5 is 70.00, which is only slightly above the top-1 of 67.68 and inconsistent with the other methods' top-5 values (86.90–87.03). These entries appear to contain typographical errors, but they also affect the comparison: the reported gains of TS, PP, and TS+PP over Baseline in these rows are partly an artifact of incorrect baseline numbers. The authors should correct these values and state whether the corrected numbers change the conclusions in Sec. 4.1.
  3. [Sec. 4.3.1 and Tables 4, 5, 9] The ablation study in Table 8 shows that attention transfer alone gives 71.51 and posterior alone gives 71.40 for 2-bit ResNet-50, while joint gives 71.96; however, Tables 4 and 9 report 'with ResNet-50' values of 71.96 and 73.25, respectively. The difference between Table 4 (joint KD only) and Table 9 (SP+KD) is 1.29 points, which is plausible, but the paper does not provide an ablation isolating the contribution of SP from the contribution of the longer 40-epoch schedule in the SP+KD setting (see the first major comment). Additionally, no sensitivity analysis is reported for β=0.5, γ=50, α_2=0.5, or the stochastic-ratio decay from 0.5, even though these parameters directly control the size of the reported gains. The authors should report at least a small grid over β and γ, or a single variation of δ, to demonstrate that the improvements are not specific to one hand-tuned configuration.
  4. [Sec. 4.3.1 and Figure 5] The claim that the jointly trained full-precision teacher can improve over the pretrained baseline is based on a comparison with the '32W, 32A' row in Table 4 (e.g., ResNet-50 teacher top-1 of 75.64 baseline versus the 'with ResNet-50' student column). However, the table reports the teacher's accuracy only in the first two rows and does not show the jointly trained teacher's accuracy after training; Figure 5 is a convergence plot without final numerical values. The statement that 'the performance of the full-precision model can be slightly improved in some cases' is therefore not directly supported by a table entry. The authors should report the final teacher accuracy for the joint-training runs, or soften the claim if the evidence is only qualitative.
  5. [Table 7] Table 7 reports results for the fixed-teacher setting without standard deviations, even though Tables 4–6 report three repeated runs with standard deviations. Since the comparison between fixed and joint teachers is used to support the claim that joint adaptation is better, the absence of repeated runs makes it impossible to assess whether the differences (e.g., 65.09 vs. 65.67 for 2-bit PreResNet-18) are statistically meaningful. The authors should report the same repeated-run statistics for the fixed-teacher experiments.
minor comments (7)
  1. [Eq. (7)] The KL divergence in Eq. (7) is written as p_full(x_i) log(p_full(x_i)/p_low(x_i)), but Eq. (8) and Eq. (9) refer to L_KL(p_low|p_full) and L_KL(p_full|p_low), respectively, which is inconsistent with the direction in Eq. (7). The authors should make the argument order explicit and consistent throughout Sec. 3.4.
  2. [Algorithm 3] Algorithm 3 uses δ_t as both a probability and a linearly decayed value, with step 12 updating δ_{t+1} = δ_t − μ; however, the text states that δ is 'linearly decreased to 0 at the 20-th epoch.' The relationship between μ and the epoch count is not defined, so the reader cannot reproduce the exact decay schedule. A sentence defining μ in terms of epochs would suffice.
  3. [Table 3] Table 3 reports results with one decimal place (e.g., 64.8, 85.7) for GroupNet baselines, while Table 2 reports the same or similar numbers with two decimal places. This makes direct comparison across tables awkward and suggests different runs or rounding conventions.
  4. [Sec. 4.3.2] The initial learning rate of 0.1 for the from-scratch student in Sec. 4.3.2 is very high compared with the 0.005 used in fine-tuning; the authors do not comment on whether the comparison between 'from scratch' and 'fine-tuning' is affected by this difference in learning rate, as opposed to the initialization alone.
  5. [Sec. 4.5] The sentence 'we find a 3.06% relative gap between the baseline on ResNet-50' is unclear: the numbers in Table 9 show a 3.06-point absolute difference (73.25 vs. 70.19), not a relative gap. The wording should distinguish absolute from relative improvement.
  6. [Figure 2] Figure 2 legend refers to 'stage-1' and 'stage-2' but does not explicitly define which stage corresponds to which quantization configuration; the caption should briefly restate that stage-1 quantizes weights only and stage-2 quantizes both weights and activations.
  7. [Sec. 4.3.4] Table 8 is labeled an 'Ablation study' but the spelling 'Abalation' appears in the caption; also, the table only reports ResNet-50 results, so the claim that 'integrating both posterior-based and attention-based distillation strategies achieve the best result' is only verified for one architecture.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity is found: the paper's central claims are empirical accuracy comparisons against external benchmarks, not derivations from fitted inputs.

full rationale

The paper proposes three training heuristics—progressive quantization, stochastic precision, and joint knowledge distillation—and supports them with experiments on ImageNet and CIFAR-100. The equations defining the method (Eqs. (1)-(9)) are the DoReFa quantizer, the straight-through estimator, and distillation losses; none is fitted to the reported accuracies or is equated to the claimed outcome. The reported gains, such as the 2-bit ResNet-50 improvement from 70.19 baseline to 72.23 with SP and 73.25 with SP+KD, are benchmark measurements evaluated against the DoReFa-Net baseline, an external method, and also against LQ-Net, BiReal-Net, and GroupNet. The paper does extend the authors' own prior conference paper [23] and uses their own GroupNet [33] as a comparison baseline, but these self-citations are descriptive rather than load-bearing: no uniqueness theorem or prior derivation is invoked to force the method's choices. The skeptic's observation that SP experiments use a 40-epoch schedule with later learning-rate decays while the stated default for baselines is 30 epochs (Sec. 4.2 vs. Sec. 4) is a potential experimental-fairness and validity concern, not circularity, because it does not make the claimed improvement equivalent to the method's inputs by construction. No fitted parameter is renamed as a prediction, and no equation reduces to itself. Thus the paper is self-contained against external benchmarks with only minor, non-load-bearing self-citation.

Assumptions & free parameters 6 free parameters · 3 assumptions · 0 invented entities

The paper contributes training heuristics rather than a theory; it introduces no new mathematical entities. The central numbers rest on hand-picked hyperparameters and the STE approximation, which are standard in the field but underexplored here.

free parameters (6)
  • stochastic ratio delta = 0.5, linearly decayed to 0 at epoch 20
    Controls the fraction of fragments quantized at each iteration; chosen by hand with no sensitivity study.
  • attention loss weight gamma = 50
    Balance between classification and attention-matching losses in joint distillation; set without sensitivity analysis.
  • KL distillation weight beta = 0.5
    Weight for posterior-matching loss in joint distillation.
  • classification loss weights alpha1, alpha2 = 1 and 0.5
    Weights for teacher and student cross-entropy losses.
  • learning rate schedule = 0.005 (student), 0.001 (teacher), decay 10x at epochs
    Standard optimization hyperparameters, not tied to a stated theory.
  • precision descent sequence = 32, 8, 4, 2 bits
    Used in progressive precision; other schedules are not explored.
assumptions (3)
  • domain assumption Straight-through estimator gradient approximation (Eq. 4) is an adequate surrogate for the true quantizer gradient.
    All three methods rely on STE to train through non-differentiable quantization; no theoretical justification is given.
  • domain assumption A pretrained full-precision model provides a useful initialization and guidance signal for low-precision training.
    Implementation details state that the authors fine-tune from the pretrained full-precision model; most experiments depend on this.
  • domain assumption Quantizing the teacher feature maps (Eq. 6) makes teacher and student activations comparable for attention transfer.
    Equation (6) measures distance between normalized quantized teacher attention maps and student attention maps; no empirical validation that this is the right correspondence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations." pith.science (2026). https://pith.science/paper/SXMJKKYO

@misc{pith2026190804680,
  author       = {Pith},
  title        = {Pith review of: Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SXMJKKYO}},
  note         = {Machine review of arXiv:1908.04680}
}
read the original abstract

This paper tackles the problem of training a deep convolutional neural network of both low-bitwidth weights and activations. Optimizing a low-precision network is very challenging due to the non-differentiability of the quantizer, which may result in substantial accuracy loss. To address this, we propose three practical approaches, including (i) progressive quantization; (ii) stochastic precision; and (iii) joint knowledge distillation to improve the network training. First, for progressive quantization, we propose two schemes to progressively find good local minima. Specifically, we propose to first optimize a net with quantized weights and subsequently quantize activations. This is in contrast to the traditional methods which optimize them simultaneously. Furthermore, we propose a second progressive quantization scheme which gradually decreases the bit-width from high-precision to low-precision during training. Second, to alleviate the excessive training burden due to the multi-round training stages, we further propose a one-stage stochastic precision strategy to randomly sample and quantize sub-networks while keeping other parts in full-precision. Finally, we adopt a novel learning scheme to jointly train a full-precision model alongside the low-precision one. By doing so, the full-precision model provides hints to guide the low-precision model training and significantly improves the performance of the low-precision network. Extensive experiments on various datasets (e.g., CIFAR-100, ImageNet) show the effectiveness of the proposed methods.

Figures

Figures reproduced from arXiv: 1908.04680 by the authors.

Figure 1
Figure 1. Demonstration of the guided training strategy. Dashed lines show the guidance loss. Similar to [19], we can also employ the posterior prob￾ability as the guidance signal. Let pfull and plow be the full-precision teacher network and low-precision student network predictions, respectively. To measure the correla￾tion between the two distributions, we employ the Kull￾back–Leibler (KL) divergence: LKL(pfull|plow) = X N … view at source ↗
Figure 3
Figure 3. The progressive training approach on AlexNet*. 4.2 Effect of the stochastic precision In this subsection, we further explore the effect of the stochastic precision strategy on general quantization ap￾proaches. The stochastic ratio δ is initialized to 0.5 and lin￾early decayed to 0 at the 20-th epoch. We train a maximum 40 epochs and decay the learning rate by 10 at the 25-th and 35-th epochs. Other hyperparameters a… view at source ↗
Figure 2
Figure 2. The two-stage training approach on ResNet-50. we apply AlexNet and ResNet-50 on the ImageNet dataset. We continuously quantize both weights and activations si￾multaneously from 32-bit→8-bit→4-bit→2-bit and explicitly illustrate the accuracy change process for each precision in [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The stochastic precision training approach on ResNet￾50. layerdrop and blockdrop respectively. We further incorporate the randomness of quantizing weights and activations into m and is denoted by W/A. From the results, we observe that all the four cases show improved p…
Figure 6
Figure 6. Figure 6: Student is learnt from scratch while teacher is fine￾tuned. PreResNet-50 is used here. 4.3.3 Learning from the fixed teacher In this section, we fix the pretrained teacher network and only fine-tune the student network. This is the scheme used by [19] to train their st…
Figure 7
Figure 7. Figure 7: Joint training vs. fixed teacher using AlexNet* on ImageNet. 4.3.4 Ablation study on guidance signals In this part, we further explore the effect of different distil￾lation guidance signals as introduced in Sec. 3.4. The results are reported in [PITH_FULL_IMAGE:figure…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

78 extracted references · 70 canonical work pages

  1. [1]

    Imagenet classi- fication with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classi- fication with deep convolutional neural networks,” in Proc. Adv. Neural Inf. Process. Syst., 2012, pp. 1097–1105. 1, 6

  2. [2]

    Very deep convolutional net- works for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional net- works for large-scale image recognition,” in Proc. Int. Conf. Learn. Repren., 2015. 1

  3. [3]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2016, pp. 770–778. 1, 6

  4. [4]

    Discrimination-aware channel pruning for deep neural networks,

    Z. Zhuang, M. Tan, B. Zhuang, J. Liu, Y. Guo, Q. Wu, J. Huang, and J. Zhu, “Discrimination-aware channel pruning for deep neural networks,” in Proc. Adv. Neural Inf. Process. Syst. , 2018, pp. 875–

  5. [5]

    Channel pruning for accelerating very deep neural networks,

    Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” in Proc. IEEE Int. Conf. Comp. Vis. , vol. 2, 2017, p. 6. 1

  6. [6]

    Pruning filters for efficient convnets,

    H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P . Graf, “Pruning filters for efficient convnets,” in Proc. Adv. Neural Inf. Process. Syst.,

  7. [7]

    Com- pression of deep convolutional neural networks for fast and low power mobile applications,

    Y.-D. Kim, E. Park, S. Yoo, T. Choi, L. Yang, and D. Shin, “Com- pression of deep convolutional neural networks for fast and low power mobile applications,” arXiv preprint arXiv:1511.06530, 2015. 1

  8. [8]

    Accelerating very deep convolutional networks for classification and detection,

    X. Zhang, J. Zou, K. He, and J. Sun, “Accelerating very deep convolutional networks for classification and detection,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 38, no. 10, pp. 1943–1955,

Show all 78 references
  1. [9]

    Incremental network quantization: Towards lossless cnns with low-precision weights,

    A. Zhou, A. Yao, Y. Guo, L. Xu, and Y. Chen, “Incremental network quantization: Towards lossless cnns with low-precision weights,” in Proc. Int. Conf. Learn. Repren., 2017. 1, 3, 4, 5

  2. [10]

    Binaryconnect: Train- ing deep neural networks with binary weights during propaga- tions,

    M. Courbariaux, Y. Bengio, and J.-P . David, “Binaryconnect: Train- ing deep neural networks with binary weights during propaga- tions,” in Proc. Adv. Neural Inf. Process. Syst. , 2015, pp. 3123–3131. 1

  3. [11]

    Trained ternary quanti- zation,

    C. Zhu, S. Han, H. Mao, and W. J. Dally, “Trained ternary quanti- zation,” in Proc. Int. Conf. Learn. Repren., 2017. 1, 6

  4. [12]

    Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients,

    S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou, “Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients,” arXiv preprint arXiv:1606.06160 , 2016. 1, 2, 3, 6

  5. [13]

    Single path one-shot neural architecture search with uniform sampling,

    Z. Guo, X. Zhang, H. Mu, W. Heng, Z. Liu, Y. Wei, and J. Sun, “Single path one-shot neural architecture search with uniform sampling,” arXiv preprint arXiv:1904.00420, 2019. 1, 2

  6. [14]

    Mobilenets: Efficient convolutional neural networks for mobile vision applications,

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017. 1, 2

  7. [15]

    Shufflenet: An extremely efficient convolutional neural network for mobile devices,

    X. Zhang, X. Zhou, M. Lin, and J. Sun, “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” inProc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018, pp. 6848–6856. 1, 2

  8. [16]

    Dropout: a simple way to prevent neural networks from overfitting,

    N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, no. 1, pp. 1929–1958, 2014. 1, 3

  9. [17]

    Deep networks with stochastic depth,

    G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2016, pp. 646–661. 1, 3

  10. [18]

    Fitnets: Hints for thin deep nets,

    A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for thin deep nets,” Proc. Int. Conf. Learn. Repren., 2015. 1, 2, 3, 5

  11. [19]

    Distilling the knowledge in a neural network,

    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in Proc. Adv. Neural Inf. Process. Syst. Workshops ,

  12. [20]

    Actor-mimic: Deep multitask and transfer reinforcement learning,

    E. Parisotto, J. L. Ba, and R. Salakhutdinov, “Actor-mimic: Deep multitask and transfer reinforcement learning,” Proc. Int. Conf. Learn. Repren., 2016. 1, 5

  13. [21]

    Paying more attention to atten- tion: Improving the performance of convolutional neural networks via attention transfer,

    S. Zagoruyko and N. Komodakis, “Paying more attention to atten- tion: Improving the performance of convolutional neural networks via attention transfer,” in Proc. Int. Conf. Learn. Repren. , 2017. 1, 2, 3, 5

  14. [22]

    Do deep nets really need to be deep?

    J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Proc. Adv. Neural Inf. Process. Syst., 2014, pp. 2654–2662. 1, 5

  15. [23]

    Towards effective low-bitwidth convolutional neural networks,

    B. Zhuang, C. Shen, M. Tan, L. Liu, and I. Reid, “Towards effective low-bitwidth convolutional neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018. 2, 3, 6 12 TABLE 9: Accuracy (%) of ResNets on the ImageNet validation set using SP and KD. Experiments are repe...

  16. [24]

    Xnor- net: Imagenet classification using binary convolutional neural networks,

    M. Rastegari, V . Ordonez, J. Redmon, and A. Farhadi, “Xnor- net: Imagenet classification using binary convolutional neural networks,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2016, pp. 525–

  17. [25]

    Binarized neural networks,

    I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized neural networks,” in Proc. Adv. Neural Inf. Process. Syst., 2016, pp. 4107–4115. 2, 3

  18. [26]

    Training competitive binary neural networks from scratch,

    J. Bethge, M. Bornstein, A. Loy, H. Yang, and C. Meinel, “Training competitive binary neural networks from scratch,” arXiv preprint arXiv:1812.01965, 2018. 2

  19. [27]

    Learning to train a binary neural network,

    J. Bethge, H. Yang, C. Bartz, and C. Meinel, “Learning to train a binary neural network,” arXiv preprint arXiv:1809.10463, 2018. 2

  20. [28]

    How to train a compact binary neural network with high accuracy?

    W. Tang, G. Hua, and L. Wang, “How to train a compact binary neural network with high accuracy?” in Proc. Twenty-Eighth AAAI Conf. on Arti. Intel., 2017, pp. 2625–2631. 2

  21. [29]

    Network sketching: Exploiting binary structure in deep cnns,

    Y. Guo, A. Yao, H. Zhao, and Y. Chen, “Network sketching: Exploiting binary structure in deep cnns,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 5955–5963. 2

  22. [30]

    Bi- real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm,

    Z. Liu, B. Wu, W. Luo, X. Yang, W. Liu, and K.-T. Cheng, “Bi- real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2, 6, 8

  23. [31]

    Pact: Parameterized clipping activation for quantized neural networks,

    J. Choi, Z. Wang, S. Venkataramani, P . I.-J. Chuang, V . Srinivasan, and K. Gopalakrishnan, “Pact: Parameterized clipping activation for quantized neural networks,” arXiv preprint arXiv:1805.06085 ,

  24. [32]

    Learned step size quantization,

    S. K. Esser, J. L. McKinstry, D. Bablani, R. Appuswamy, and D. S. Modha, “Learned step size quantization,” in Proc. Int. Conf. Learn. Repren., 2020. 2

  25. [33]

    Strutured binary neural network for accurate image classification and semantic segmentation,

    B. Zhuang, C. Shen, M. Tan, L. Liu, and I. Reid, “Strutured binary neural network for accurate image classification and semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019. 2, 6

  26. [34]

    Towards accurate binary convolu- tional neural network,

    X. Lin, C. Zhao, and W. Pan, “Towards accurate binary convolu- tional neural network,” in Proc. Adv. Neural Inf. Process. Syst., 2017, pp. 344–352. 2

  27. [35]

    Deep learning with low precision by half-wave gaussian quantization,

    Z. Cai, X. He, J. Sun, and N. Vasconcelos, “Deep learning with low precision by half-wave gaussian quantization,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 5918–5926. 2, 6

  28. [36]

    Lq-nets: Learned quanti- zation for highly accurate and compact deep neural networks,

    D. Zhang, J. Yang, D. Ye, and G. Hua, “Lq-nets: Learned quanti- zation for highly accurate and compact deep neural networks,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2, 6

  29. [37]

    Learning to quantize deep networks by optimizing quantization intervals with task loss,

    S. Jung, C. Son, S. Lee, J. Son, J.-J. Han, Y. Kwak, S. J. Hwang, and C. Choi, “Learning to quantize deep networks by optimizing quantization intervals with task loss,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 4350–4359. 2

  30. [38]

    Estimating or propagat- ing gradients through stochastic neurons for conditional compu- tation,

    Y. Bengio, N. L ´eonard, and A. Courville, “Estimating or propagat- ing gradients through stochastic neurons for conditional compu- tation,” arXiv preprint arXiv:1308.3432, 2013. 2, 3

  31. [39]

    Loss-aware weight quantization of deep networks,

    L. Hou and J. T. Kwok, “Loss-aware weight quantization of deep networks,” in Proc. Int. Conf. Learn. Repren., 2018. 2

  32. [40]

    Regularizing activation distribution for training binarized deep networks,

    R. Ding, T.-W. Chin, Z. Liu, and D. Marculescu, “Regularizing activation distribution for training binarized deep networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 11 408–11 417. 2

  33. [41]

    Learning low preci- sion deep neural networks through regularization,

    Y. Choi, M. El-Khamy, and J. Lee, “Learning low preci- sion deep neural networks through regularization,” ArXiv, vol. abs/1809.00095, 2018. 2

  34. [42]

    Proxquant: Quantized neural networks via proximal operators,

    Y. Bai, Y.-X. Wang, and E. Liberty, “Proxquant: Quantized neural networks via proximal operators,” in Proc. Int. Conf. Learn. Repren.,

  35. [43]

    Weighted-entropy-based quantization for deep neural networks,

    E. Park, J. Ahn, and S. Yoo, “Weighted-entropy-based quantization for deep neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 5456–5464. 2 13

  36. [44]

    Model compression via distillation and quantization,

    A. Polino, R. Pascanu, and D. Alistarh, “Model compression via distillation and quantization,” in Proc. Int. Conf. Learn. Repren. ,

  37. [45]

    Relaxed quantization for discretized neural networks,

    C. Louizos, M. Reisser, T. Blankevoort, E. Gavves, and M. Welling, “Relaxed quantization for discretized neural networks,” in Proc. Int. Conf. Learn. Repren., 2019. 2

  38. [46]

    Ai benchmark: Running deep neural networks on android smartphones,

    A. Ignatov, R. Timofte, W. Chou, K. Wang, M. Wu, T. Hartley, and L. Van Gool, “Ai benchmark: Running deep neural networks on android smartphones,” in Proc. Eur. Conf. Comp. Vis., 2018. 2

  39. [47]

    Bmxnet: An open- source binary neural network implementation based on mxnet,

    H. Yang, M. Fritzsche, C. Bartz, and C. Meinel, “Bmxnet: An open- source binary neural network implementation based on mxnet,” in Proc. of the ACM Int. Conf. on Multimedia. ACM, 2017, pp. 1209–

  40. [48]

    Finn: A framework for fast, scalable binarized neural network inference,

    Y. Umuroglu, N. J. Fraser, G. Gambardella, M. Blott, P . Leong, M. Jahre, and K. Vissers, “Finn: A framework for fast, scalable binarized neural network inference,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. ACM, 2017, pp. 65–74. 2

  41. [49]

    Quantization and training of neural networks for efficient integer-arithmetic-only inference,

    B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018. 2

  42. [50]

    Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size,

    F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size,” arXiv preprint arXiv:1602.07360, 2016. 2

  43. [51]

    Xception: Deep learning with depthwise separable convolutions,

    F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 1251–1258. 2

  44. [52]

    Neural architecture search with reinforce- ment learning,

    B. Zoph and Q. V . Le, “Neural architecture search with reinforce- ment learning,” in Proc. Int. Conf. Learn. Repren., 2017. 2

  45. [53]

    Efficient neural architecture search via parameter sharing,

    H. Pham, M. Y. Guan, B. Zoph, Q. V . Le, and J. Dean, “Efficient neural architecture search via parameter sharing,” in Proc. Int. Conf. Mach. Learn., 2018. 2

  46. [54]

    Learning trans- ferable architectures for scalable image recognition,

    B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning trans- ferable architectures for scalable image recognition,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018. 2

  47. [55]

    Progressive neural architecture search,

    C. Liu, B. Zoph, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2

  48. [56]

    Regularized evolution for image classifier architecture search,

    E. Real, A. Aggarwal, Y. Huang, and Q. V . Le, “Regularized evolution for image classifier architecture search,” in Proceedings of the aaai conference on artificial intelligence , vol. 33, 2019, pp. 4780–

  49. [57]

    Darts: Differentiable architec- ture search,

    H. Liu, K. Simonyan, and Y. Yang, “Darts: Differentiable architec- ture search,” in Proc. Int. Conf. Learn. Repren., 2019. 2

  50. [58]

    Proxylessnas: Direct neural architec- ture search on target task and hardware,

    H. Cai, L. Zhu, and S. Han, “Proxylessnas: Direct neural architec- ture search on target task and hardware,” in Proc. Int. Conf. Learn. Repren., 2019. 2

  51. [59]

    Nisp: Pruning networks using neuron importance score propagation,

    R. Yu, A. Li, C.-F. Chen, J.-H. Lai, V . I. Morariu, X. Han, M. Gao, C.-Y. Lin, and L. S. Davis, “Nisp: Pruning networks using neuron importance score propagation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018, pp. 9194–9203. 2

  52. [60]

    N2n learning: Network to network compression via policy gradient reinforcement learning,

    A. Ashok, N. Rhinehart, F. Beainy, and K. M. Kitani, “N2n learning: Network to network compression via policy gradient reinforcement learning,” in Proc. Int. Conf. Learn. Repren., 2018. 2

  53. [61]

    Amc: Automl for model compression and acceleration on mobile devices,

    Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “Amc: Automl for model compression and acceleration on mobile devices,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2

  54. [62]

    Clip-q: Deep network compression learning by in-parallel pruning-quantization,

    F. Tung and G. Mori, “Clip-q: Deep network compression learning by in-parallel pruning-quantization,” inProc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018, pp. 7873–7882. 2

  55. [63]

    Network pruning via transformable archi- tecture search,

    X. Dong and Y. Yang, “Network pruning via transformable archi- tecture search,” in Proc. Adv. Neural Inf. Process. Syst. , 2019, pp. 760–771. 2

  56. [64]

    Real-time action recognition with enhanced motion vector cnns,

    B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang, “Real-time action recognition with enhanced motion vector cnns,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2016, pp. 2718–2726. 3

  57. [65]

    Learning efficient object detection models with knowledge distillation,

    G. Chen, W. Choi, X. Yu, T. Han, and M. Chandraker, “Learning efficient object detection models with knowledge distillation,” in Proc. Adv. Neural Inf. Process. Syst., 2017, pp. 742–751. 3

  58. [66]

    Quantization mimic: Towards very tiny cnn for object detection,

    Y. Wei, X. Pan, H. Qin, W. Ouyang, and J. Yan, “Quantization mimic: Towards very tiny cnn for object detection,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018, pp. 267–283. 3

  59. [67]

    Knowledge adaptation for efficient semantic segmentation,

    T. He, C. Shen, Z. Tian, D. Gong, C. Sun, and Y. Yan, “Knowledge adaptation for efficient semantic segmentation,” arXiv preprint arXiv:1903.04688, 2019. 3

  60. [68]

    Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,

    A. Mishra and D. Marr, “Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,” in Proc. Int. Conf. Learn. Repren., 2018. 3

  61. [69]

    Maxout networks,

    I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, “Maxout networks,” in Proc. Int. Conf. Mach. Learn. ,

  62. [70]

    Regular- ization of neural networks using dropconnect,

    L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus, “Regular- ization of neural networks using dropconnect,” in Proc. Int. Conf. Mach. Learn., 2013, pp. 1058–1066. 3

  63. [71]

    Gradual dropin of layers to train very deep neural networks,

    L. N. Smith, E. M. Hand, and T. Doster, “Gradual dropin of layers to train very deep neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2016, pp. 4763–4771. 3

  64. [72]

    Learning accurate low-bit deep neural networks with stochastic quantization,

    Y. Dong, R. Ni, J. Li, Y. Chen, J. Zhu, and H. Su, “Learning accurate low-bit deep neural networks with stochastic quantization,” in Proc. Brit. Mach. Vis. Conf., 2017. 3

  65. [73]

    Slimmable neural networks,

    J. Yu, L. Yang, N. Xu, J. Yang, and T. Huang, “Slimmable neural networks,” 2019. 3, 4

  66. [74]

    Universally slimmable networks and improved training techniques,

    J. Yu and T. S. Huang, “Universally slimmable networks and improved training techniques,” in Proc. IEEE Int. Conf. Comp. Vis., 2019, pp. 1803–1811. 3

  67. [75]

    Learning multiple layers of fea- tures from tiny images,

    A. Krizhevsky and G. Hinton, “Learning multiple layers of fea- tures from tiny images,” 2009. 6

  68. [76]

    Imagenet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” Int. J. Comp. Vis., vol. 115, no. 3, pp. 211–252, 2015. 6

  69. [77]

    Identity mappings in deep residual networks,

    K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2016, pp. 630–645. 6

  70. [78]

    Precision highway for ultra low-precision quantization,

    E. Park, D. Kim, S. Yoo, and P . Vajda, “Precision highway for ultra low-precision quantization,” arXiv preprint arXiv:1812.09818, 2018. 8 Bohan Zhuang is a Lecturer at Monash University, Australia. He re- ceived his PhD degree from The University of Adelaide. Mingkui Tan is a...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.