REVIEW 5 major objections 7 minor 78 references
Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
T0 review · 5 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Low-bit convolutional networks can be trained to beat their quantized baselines.
desk verdict Useful quantization training tricks with a real confound: SP gains may be training-budget artifacts, but KD and progressive quantization hold up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the training schedule itself, built on three mechanisms. First, a fixed-point quantizer of the DoReFa form, $Q(z) = \frac{1}{2^k-1}\mathrm{round}((2^k-1)z_r)$, is combined with a straight-through gradient approximation, $\partial z_q/\partial z_r \approx 1$, so gradient-based updates can flow through a non-differentiable quantizer. Second, a binary mask over network fragments (layers or residual blocks, and optionally weights versus activations) selects which parts are quantized at each iteration, with the stochastic ratio $\delta$ starting at 0.5 and linearly decaying to zero; this provides the progressive-relaxation effect in one stage. Third, a two-headed loss couples the low-precision student to a jointly updated full-precision teacher through KL divergence on posterior probabilities, weighted by $\beta = 0.5$, and attention transfer on feature maps, weighted by $\gamma = 50$. These mechanisms carry the argument by making the discrete optimization tractable and by continuously guiding the student toward the teacher during training.
What would settle it
Re-run the 2-bit ResNet-50 ImageNet experiment with the paper's exact hyperparameters but disable stochastic precision while keeping joint distillation: the paper reports 71.96% top-1 without SP and 73.25% with it, so a replication showing no statistically significant gap would falsify the claim that stochastic precision adds a real benefit; similarly, disabling the distillation losses should reproduce the 70.19% baseline if the teacher guidance is genuinely responsible for the gain.
Extended reading notes
Core claim
The paper's central claim is that the optimization difficulty of quantized neural networks, not the capacity of the low-precision representation, is the main obstacle to accurate low-bit models, and that this difficulty can be attacked with three training techniques. Progressive quantization first solves the easier problem of quantized weights with full-precision activations, then adds activation quantization, and also anneals the bit-width from 32 bits down to the target precision. Stochastic precision randomly quantizes a fraction of layers, blocks, weights, or activations per iteration while the stochastic ratio is linearly decayed to zero, easing gradient flow in a single training stage. Joint knowledge distillation trains a full-precision teacher and a low-precision student together, using both posterior KL divergence and attention-transfer losses, and reports that the student surpasses the baseline and sometimes the teacher improves as well. The paper reports that all low-precision models produced by these methods surpass their corresponding baselines, with the largest combined gain reaching 73.25% top-1 accuracy for 2-bit ResNet-50 on ImageNet.
Load-bearing premise
The load-bearing premise is that the straight-through estimator supplies genuinely useful gradient directions at very low precision, and that the hand-set schedule of stochastic ratios and distillation weights keeps working for other architectures, datasets, and quantization schemes.
Editorial extensions
If this is right
- 2-bit ResNet-50 on ImageNet can reach 73.25% top-1 accuracy with stochastic precision plus joint distillation, compared with 70.19% for the DoReFa-Net baseline, closing most of the gap to the 75.64% full-precision model.
- The three training strategies are complementary: two-stage optimization, progressive precision, stochastic precision, and knowledge distillation can be combined, and each combination in the paper improves over the corresponding baseline.
- Joint distillation helps more when the low-precision student is harder to train, since learning from scratch shows larger relative gains than fine-tuning, and quantized ResNet benefits more than quantized PreResNet.
- The training techniques are orthogonal to the quantizer: stochastic precision improves DoReFa-Net, LQ-Net, BiReal-Net, and GroupNet baselines, suggesting the methods could be transferred to other quantization designs.
- The full-precision teacher can also be slightly improved by joint training with the student, indicating that the mutual adaptation acts as a regularizer rather than a one-way transfer.
Reading between the lines
- A natural extension, not pursued in the paper, would be to replace the hand-set stochastic ratio decay and distillation weights with learned or per-layer schedules, since the paper gives no sensitivity analysis for these hyperparameters and the gains might be further improved or made more robust.
- Beyond image classification, one could test the same three training strategies on low-bit object detection or semantic segmentation, where the optimization difficulty is typically greater and the reported benefits might be even more pronounced.
- The observation that the jointly updated teacher improves suggests that mutual distillation could serve as a general regularizer for full-precision training, independent of quantization.
- Because the methods are quantizer-agnostic, a testable extension is to pair them with modern learned-step-size quantizers or with mixed-precision allocation, potentially recovering additional accuracy at the same average bit-width.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes three training strategies for low-bitwidth CNNs: (i) progressive quantization, with a two-stage scheme (quantize weights first, then activations) and a progressive precision scheme (32→8→4→2 bits during training); (ii) stochastic precision (SP), which randomly quantizes fragments of the network while keeping the rest full-precision, with the stochastic ratio decaying to zero; and (iii) joint knowledge distillation (KD), in which a full-precision teacher and a low-precision student are trained together using both KL-divergence and attention-transfer losses. Experiments on CIFAR-100 and ImageNet with DoReFa-Net, LQ-Net, BiReal-Net, and GroupNet report consistent accuracy improvements over the corresponding baselines, culminating in combined SP+KD results such as 73.25% top-1 for 2-bit ResNet-50 versus a 70.19% baseline.
Significance. If the reported gains hold under controlled comparisons, the paper would make a useful practical contribution: it offers simple, quantizer-agnostic training recipes that improve low-bitword accuracy across several architectures and quantization schemes, and the combination of stochastic precision with joint distillation appears complementary rather than redundant. The paper also ships a fairly large set of experiments with repeated runs and standard deviations for most of the headline tables, which is a strength. However, the central quantitative claim is weakened by a training-budget confound in the SP and SP+KD experiments, by apparent typographical inconsistencies in Table 1, and by the absence of sensitivity analysis for the key hyperparameters (δ, β, γ, and the precision schedule). The contribution is incremental over the authors' CVPR 2018 conference paper, but the new one-stage SP formulation and the joint-teacher distillation study constitute a reasonable extension for a journal version.
major comments (5)
- [Sec. 4.2 and Table 2] The SP experiments use a 40-epoch schedule (learning-rate decays at epochs 25 and 35), while the baselines are trained for 30 epochs with decays at epochs 15 and 25, as stated in Sec. 4. The baseline numbers in Table 2 and Table 9 are identical to those in Table 4, which uses the default 30-epoch schedule. No baseline re-run under the 40-epoch schedule is reported, so the gains attributed to SP (e.g., 70.19→72.23 in Table 2) and to SP+KD (70.19→73.25 in Table 9) are confounded by the additional 10 epochs of training and the later learning-rate decays. The authors should either re-run the baselines under the SP schedule or provide evidence, such as a 40-epoch baseline curve, that the extra epochs alone do not explain the improvement. Without this, the central claim that 'all our low-precision models surpass the corresponding baselines' is not established on a like-for-like training budget.
- [Table 1] The top-5 accuracy entries for the 4W,4A ResNet-50* row are internally inconsistent: the Baseline is listed as 75.70, which is below the Top-1 value of 75.11, and is much lower than the Baseline+TS value of 91.93. Similarly, the 2W,2A ResNet-50* Baseline top-5 is 70.00, which is only slightly above the top-1 of 67.68 and inconsistent with the other methods' top-5 values (86.90–87.03). These entries appear to contain typographical errors, but they also affect the comparison: the reported gains of TS, PP, and TS+PP over Baseline in these rows are partly an artifact of incorrect baseline numbers. The authors should correct these values and state whether the corrected numbers change the conclusions in Sec. 4.1.
- [Sec. 4.3.1 and Tables 4, 5, 9] The ablation study in Table 8 shows that attention transfer alone gives 71.51 and posterior alone gives 71.40 for 2-bit ResNet-50, while joint gives 71.96; however, Tables 4 and 9 report 'with ResNet-50' values of 71.96 and 73.25, respectively. The difference between Table 4 (joint KD only) and Table 9 (SP+KD) is 1.29 points, which is plausible, but the paper does not provide an ablation isolating the contribution of SP from the contribution of the longer 40-epoch schedule in the SP+KD setting (see the first major comment). Additionally, no sensitivity analysis is reported for β=0.5, γ=50, α_2=0.5, or the stochastic-ratio decay from 0.5, even though these parameters directly control the size of the reported gains. The authors should report at least a small grid over β and γ, or a single variation of δ, to demonstrate that the improvements are not specific to one hand-tuned configuration.
- [Sec. 4.3.1 and Figure 5] The claim that the jointly trained full-precision teacher can improve over the pretrained baseline is based on a comparison with the '32W, 32A' row in Table 4 (e.g., ResNet-50 teacher top-1 of 75.64 baseline versus the 'with ResNet-50' student column). However, the table reports the teacher's accuracy only in the first two rows and does not show the jointly trained teacher's accuracy after training; Figure 5 is a convergence plot without final numerical values. The statement that 'the performance of the full-precision model can be slightly improved in some cases' is therefore not directly supported by a table entry. The authors should report the final teacher accuracy for the joint-training runs, or soften the claim if the evidence is only qualitative.
- [Table 7] Table 7 reports results for the fixed-teacher setting without standard deviations, even though Tables 4–6 report three repeated runs with standard deviations. Since the comparison between fixed and joint teachers is used to support the claim that joint adaptation is better, the absence of repeated runs makes it impossible to assess whether the differences (e.g., 65.09 vs. 65.67 for 2-bit PreResNet-18) are statistically meaningful. The authors should report the same repeated-run statistics for the fixed-teacher experiments.
minor comments (7)
- [Eq. (7)] The KL divergence in Eq. (7) is written as p_full(x_i) log(p_full(x_i)/p_low(x_i)), but Eq. (8) and Eq. (9) refer to L_KL(p_low|p_full) and L_KL(p_full|p_low), respectively, which is inconsistent with the direction in Eq. (7). The authors should make the argument order explicit and consistent throughout Sec. 3.4.
- [Algorithm 3] Algorithm 3 uses δ_t as both a probability and a linearly decayed value, with step 12 updating δ_{t+1} = δ_t − μ; however, the text states that δ is 'linearly decreased to 0 at the 20-th epoch.' The relationship between μ and the epoch count is not defined, so the reader cannot reproduce the exact decay schedule. A sentence defining μ in terms of epochs would suffice.
- [Table 3] Table 3 reports results with one decimal place (e.g., 64.8, 85.7) for GroupNet baselines, while Table 2 reports the same or similar numbers with two decimal places. This makes direct comparison across tables awkward and suggests different runs or rounding conventions.
- [Sec. 4.3.2] The initial learning rate of 0.1 for the from-scratch student in Sec. 4.3.2 is very high compared with the 0.005 used in fine-tuning; the authors do not comment on whether the comparison between 'from scratch' and 'fine-tuning' is affected by this difference in learning rate, as opposed to the initialization alone.
- [Sec. 4.5] The sentence 'we find a 3.06% relative gap between the baseline on ResNet-50' is unclear: the numbers in Table 9 show a 3.06-point absolute difference (73.25 vs. 70.19), not a relative gap. The wording should distinguish absolute from relative improvement.
- [Figure 2] Figure 2 legend refers to 'stage-1' and 'stage-2' but does not explicitly define which stage corresponds to which quantization configuration; the caption should briefly restate that stage-1 quantizes weights only and stage-2 quantizes both weights and activations.
- [Sec. 4.3.4] Table 8 is labeled an 'Ablation study' but the spelling 'Abalation' appears in the caption; also, the table only reports ResNet-50 results, so the claim that 'integrating both posterior-based and attention-based distillation strategies achieve the best result' is only verified for one architecture.
Circularity Check
No significant circularity is found: the paper's central claims are empirical accuracy comparisons against external benchmarks, not derivations from fitted inputs.
full rationale
The paper proposes three training heuristics—progressive quantization, stochastic precision, and joint knowledge distillation—and supports them with experiments on ImageNet and CIFAR-100. The equations defining the method (Eqs. (1)-(9)) are the DoReFa quantizer, the straight-through estimator, and distillation losses; none is fitted to the reported accuracies or is equated to the claimed outcome. The reported gains, such as the 2-bit ResNet-50 improvement from 70.19 baseline to 72.23 with SP and 73.25 with SP+KD, are benchmark measurements evaluated against the DoReFa-Net baseline, an external method, and also against LQ-Net, BiReal-Net, and GroupNet. The paper does extend the authors' own prior conference paper [23] and uses their own GroupNet [33] as a comparison baseline, but these self-citations are descriptive rather than load-bearing: no uniqueness theorem or prior derivation is invoked to force the method's choices. The skeptic's observation that SP experiments use a 40-epoch schedule with later learning-rate decays while the stated default for baselines is 30 epochs (Sec. 4.2 vs. Sec. 4) is a potential experimental-fairness and validity concern, not circularity, because it does not make the claimed improvement equivalent to the method's inputs by construction. No fitted parameter is renamed as a prediction, and no equation reduces to itself. Thus the paper is self-contained against external benchmarks with only minor, non-load-bearing self-citation.
Assumptions & free parameters
free parameters (6)
- stochastic ratio delta =
0.5, linearly decayed to 0 at epoch 20
- attention loss weight gamma =
50
- KL distillation weight beta =
0.5
- classification loss weights alpha1, alpha2 =
1 and 0.5
- learning rate schedule =
0.005 (student), 0.001 (teacher), decay 10x at epochs
- precision descent sequence =
32, 8, 4, 2 bits
assumptions (3)
- domain assumption Straight-through estimator gradient approximation (Eq. 4) is an adequate surrogate for the true quantizer gradient.
- domain assumption A pretrained full-precision model provides a useful initialization and guidance signal for low-precision training.
- domain assumption Quantizing the teacher feature maps (Eq. 6) makes teacher and student activations comparable for attention transfer.
Cite this review
Pith. "Pith review of Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations." pith.science (2026). https://pith.science/paper/SXMJKKYO
@misc{pith2026190804680,
author = {Pith},
title = {Pith review of: Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXMJKKYO}},
note = {Machine review of arXiv:1908.04680}
}
read the original abstract
This paper tackles the problem of training a deep convolutional neural network of both low-bitwidth weights and activations. Optimizing a low-precision network is very challenging due to the non-differentiability of the quantizer, which may result in substantial accuracy loss. To address this, we propose three practical approaches, including (i) progressive quantization; (ii) stochastic precision; and (iii) joint knowledge distillation to improve the network training. First, for progressive quantization, we propose two schemes to progressively find good local minima. Specifically, we propose to first optimize a net with quantized weights and subsequently quantize activations. This is in contrast to the traditional methods which optimize them simultaneously. Furthermore, we propose a second progressive quantization scheme which gradually decreases the bit-width from high-precision to low-precision during training. Second, to alleviate the excessive training burden due to the multi-round training stages, we further propose a one-stage stochastic precision strategy to randomly sample and quantize sub-networks while keeping other parts in full-precision. Finally, we adopt a novel learning scheme to jointly train a full-precision model alongside the low-precision one. By doing so, the full-precision model provides hints to guide the low-precision model training and significantly improves the performance of the low-precision network. Extensive experiments on various datasets (e.g., CIFAR-100, ImageNet) show the effectiveness of the proposed methods.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Imagenet classi- fication with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classi- fication with deep convolutional neural networks,” in Proc. Adv. Neural Inf. Process. Syst., 2012, pp. 1097–1105. 1, 6
work page 2012
-
[2]
Very deep convolutional net- works for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional net- works for large-scale image recognition,” in Proc. Int. Conf. Learn. Repren., 2015. 1
work page 2015
-
[3]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2016, pp. 770–778. 1, 6
work page 2016
-
[4]
Discrimination-aware channel pruning for deep neural networks,
Z. Zhuang, M. Tan, B. Zhuang, J. Liu, Y. Guo, Q. Wu, J. Huang, and J. Zhu, “Discrimination-aware channel pruning for deep neural networks,” in Proc. Adv. Neural Inf. Process. Syst. , 2018, pp. 875–
work page 2018
-
[5]
Channel pruning for accelerating very deep neural networks,
Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” in Proc. IEEE Int. Conf. Comp. Vis. , vol. 2, 2017, p. 6. 1
work page 2017
-
[6]
Pruning filters for efficient convnets,
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P . Graf, “Pruning filters for efficient convnets,” in Proc. Adv. Neural Inf. Process. Syst.,
-
[7]
Com- pression of deep convolutional neural networks for fast and low power mobile applications,
Y.-D. Kim, E. Park, S. Yoo, T. Choi, L. Yang, and D. Shin, “Com- pression of deep convolutional neural networks for fast and low power mobile applications,” arXiv preprint arXiv:1511.06530, 2015. 1
arXiv 2015
-
[8]
Accelerating very deep convolutional networks for classification and detection,
X. Zhang, J. Zou, K. He, and J. Sun, “Accelerating very deep convolutional networks for classification and detection,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 38, no. 10, pp. 1943–1955,
work page 1943
Show all 78 references
-
[9]
Incremental network quantization: Towards lossless cnns with low-precision weights,
A. Zhou, A. Yao, Y. Guo, L. Xu, and Y. Chen, “Incremental network quantization: Towards lossless cnns with low-precision weights,” in Proc. Int. Conf. Learn. Repren., 2017. 1, 3, 4, 5
2017
-
[10]
Binaryconnect: Train- ing deep neural networks with binary weights during propaga- tions,
M. Courbariaux, Y. Bengio, and J.-P . David, “Binaryconnect: Train- ing deep neural networks with binary weights during propaga- tions,” in Proc. Adv. Neural Inf. Process. Syst. , 2015, pp. 3123–3131. 1
2015
-
[11]
Trained ternary quanti- zation,
C. Zhu, S. Han, H. Mao, and W. J. Dally, “Trained ternary quanti- zation,” in Proc. Int. Conf. Learn. Repren., 2017. 1, 6
2017
-
[12]
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients,
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou, “Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients,” arXiv preprint arXiv:1606.06160 , 2016. 1, 2, 3, 6
2016 arXiv
-
[13]
Single path one-shot neural architecture search with uniform sampling,
Z. Guo, X. Zhang, H. Mu, W. Heng, Z. Liu, Y. Wei, and J. Sun, “Single path one-shot neural architecture search with uniform sampling,” arXiv preprint arXiv:1904.00420, 2019. 1, 2
1904 arXiv
-
[14]
Mobilenets: Efficient convolutional neural networks for mobile vision applications,
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017. 1, 2
2017 arXiv
-
[15]
Shufflenet: An extremely efficient convolutional neural network for mobile devices,
X. Zhang, X. Zhou, M. Lin, and J. Sun, “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” inProc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018, pp. 6848–6856. 1, 2
2018
-
[16]
Dropout: a simple way to prevent neural networks from overfitting,
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, no. 1, pp. 1929–1958, 2014. 1, 3
1929
-
[17]
Deep networks with stochastic depth,
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2016, pp. 646–661. 1, 3
2016
-
[18]
Fitnets: Hints for thin deep nets,
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for thin deep nets,” Proc. Int. Conf. Learn. Repren., 2015. 1, 2, 3, 5
2015
-
[19]
Distilling the knowledge in a neural network,
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in Proc. Adv. Neural Inf. Process. Syst. Workshops ,
-
[20]
Actor-mimic: Deep multitask and transfer reinforcement learning,
E. Parisotto, J. L. Ba, and R. Salakhutdinov, “Actor-mimic: Deep multitask and transfer reinforcement learning,” Proc. Int. Conf. Learn. Repren., 2016. 1, 5
2016
-
[21]
Paying more attention to atten- tion: Improving the performance of convolutional neural networks via attention transfer,
S. Zagoruyko and N. Komodakis, “Paying more attention to atten- tion: Improving the performance of convolutional neural networks via attention transfer,” in Proc. Int. Conf. Learn. Repren. , 2017. 1, 2, 3, 5
2017
-
[22]
Do deep nets really need to be deep?
J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Proc. Adv. Neural Inf. Process. Syst., 2014, pp. 2654–2662. 1, 5
2014
-
[23]
Towards effective low-bitwidth convolutional neural networks,
B. Zhuang, C. Shen, M. Tan, L. Liu, and I. Reid, “Towards effective low-bitwidth convolutional neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018. 2, 3, 6 12 TABLE 9: Accuracy (%) of ResNets on the ImageNet validation set using SP and KD. Experiments are repe...
2018
-
[24]
Xnor- net: Imagenet classification using binary convolutional neural networks,
M. Rastegari, V . Ordonez, J. Redmon, and A. Farhadi, “Xnor- net: Imagenet classification using binary convolutional neural networks,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2016, pp. 525–
2016
-
[25]
Binarized neural networks,
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized neural networks,” in Proc. Adv. Neural Inf. Process. Syst., 2016, pp. 4107–4115. 2, 3
2016
-
[26]
Training competitive binary neural networks from scratch,
J. Bethge, M. Bornstein, A. Loy, H. Yang, and C. Meinel, “Training competitive binary neural networks from scratch,” arXiv preprint arXiv:1812.01965, 2018. 2
2018 arXiv
-
[27]
Learning to train a binary neural network,
J. Bethge, H. Yang, C. Bartz, and C. Meinel, “Learning to train a binary neural network,” arXiv preprint arXiv:1809.10463, 2018. 2
2018 arXiv
-
[28]
How to train a compact binary neural network with high accuracy?
W. Tang, G. Hua, and L. Wang, “How to train a compact binary neural network with high accuracy?” in Proc. Twenty-Eighth AAAI Conf. on Arti. Intel., 2017, pp. 2625–2631. 2
2017
-
[29]
Network sketching: Exploiting binary structure in deep cnns,
Y. Guo, A. Yao, H. Zhao, and Y. Chen, “Network sketching: Exploiting binary structure in deep cnns,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 5955–5963. 2
2017
-
[30]
Bi- real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm,
Z. Liu, B. Wu, W. Luo, X. Yang, W. Liu, and K.-T. Cheng, “Bi- real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2, 6, 8
2018
-
[31]
Pact: Parameterized clipping activation for quantized neural networks,
J. Choi, Z. Wang, S. Venkataramani, P . I.-J. Chuang, V . Srinivasan, and K. Gopalakrishnan, “Pact: Parameterized clipping activation for quantized neural networks,” arXiv preprint arXiv:1805.06085 ,
-
[32]
Learned step size quantization,
S. K. Esser, J. L. McKinstry, D. Bablani, R. Appuswamy, and D. S. Modha, “Learned step size quantization,” in Proc. Int. Conf. Learn. Repren., 2020. 2
2020
-
[33]
Strutured binary neural network for accurate image classification and semantic segmentation,
B. Zhuang, C. Shen, M. Tan, L. Liu, and I. Reid, “Strutured binary neural network for accurate image classification and semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019. 2, 6
2019
-
[34]
Towards accurate binary convolu- tional neural network,
X. Lin, C. Zhao, and W. Pan, “Towards accurate binary convolu- tional neural network,” in Proc. Adv. Neural Inf. Process. Syst., 2017, pp. 344–352. 2
2017
-
[35]
Deep learning with low precision by half-wave gaussian quantization,
Z. Cai, X. He, J. Sun, and N. Vasconcelos, “Deep learning with low precision by half-wave gaussian quantization,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 5918–5926. 2, 6
2017
-
[36]
Lq-nets: Learned quanti- zation for highly accurate and compact deep neural networks,
D. Zhang, J. Yang, D. Ye, and G. Hua, “Lq-nets: Learned quanti- zation for highly accurate and compact deep neural networks,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2, 6
2018
-
[37]
Learning to quantize deep networks by optimizing quantization intervals with task loss,
S. Jung, C. Son, S. Lee, J. Son, J.-J. Han, Y. Kwak, S. J. Hwang, and C. Choi, “Learning to quantize deep networks by optimizing quantization intervals with task loss,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 4350–4359. 2
2019
-
[38]
Estimating or propagat- ing gradients through stochastic neurons for conditional compu- tation,
Y. Bengio, N. L ´eonard, and A. Courville, “Estimating or propagat- ing gradients through stochastic neurons for conditional compu- tation,” arXiv preprint arXiv:1308.3432, 2013. 2, 3
2013 arXiv
-
[39]
Loss-aware weight quantization of deep networks,
L. Hou and J. T. Kwok, “Loss-aware weight quantization of deep networks,” in Proc. Int. Conf. Learn. Repren., 2018. 2
2018
-
[40]
Regularizing activation distribution for training binarized deep networks,
R. Ding, T.-W. Chin, Z. Liu, and D. Marculescu, “Regularizing activation distribution for training binarized deep networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 11 408–11 417. 2
2019
-
[41]
Learning low preci- sion deep neural networks through regularization,
Y. Choi, M. El-Khamy, and J. Lee, “Learning low preci- sion deep neural networks through regularization,” ArXiv, vol. abs/1809.00095, 2018. 2
2018 arXiv
-
[42]
Proxquant: Quantized neural networks via proximal operators,
Y. Bai, Y.-X. Wang, and E. Liberty, “Proxquant: Quantized neural networks via proximal operators,” in Proc. Int. Conf. Learn. Repren.,
-
[43]
Weighted-entropy-based quantization for deep neural networks,
E. Park, J. Ahn, and S. Yoo, “Weighted-entropy-based quantization for deep neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 5456–5464. 2 13
2017
-
[44]
Model compression via distillation and quantization,
A. Polino, R. Pascanu, and D. Alistarh, “Model compression via distillation and quantization,” in Proc. Int. Conf. Learn. Repren. ,
-
[45]
Relaxed quantization for discretized neural networks,
C. Louizos, M. Reisser, T. Blankevoort, E. Gavves, and M. Welling, “Relaxed quantization for discretized neural networks,” in Proc. Int. Conf. Learn. Repren., 2019. 2
2019
-
[46]
Ai benchmark: Running deep neural networks on android smartphones,
A. Ignatov, R. Timofte, W. Chou, K. Wang, M. Wu, T. Hartley, and L. Van Gool, “Ai benchmark: Running deep neural networks on android smartphones,” in Proc. Eur. Conf. Comp. Vis., 2018. 2
2018
-
[47]
Bmxnet: An open- source binary neural network implementation based on mxnet,
H. Yang, M. Fritzsche, C. Bartz, and C. Meinel, “Bmxnet: An open- source binary neural network implementation based on mxnet,” in Proc. of the ACM Int. Conf. on Multimedia. ACM, 2017, pp. 1209–
2017
-
[48]
Finn: A framework for fast, scalable binarized neural network inference,
Y. Umuroglu, N. J. Fraser, G. Gambardella, M. Blott, P . Leong, M. Jahre, and K. Vissers, “Finn: A framework for fast, scalable binarized neural network inference,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. ACM, 2017, pp. 65–74. 2
2017
-
[49]
Quantization and training of neural networks for efficient integer-arithmetic-only inference,
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018. 2
2018
-
[50]
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size,
F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size,” arXiv preprint arXiv:1602.07360, 2016. 2
2016 arXiv
-
[51]
Xception: Deep learning with depthwise separable convolutions,
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 1251–1258. 2
2017
-
[52]
Neural architecture search with reinforce- ment learning,
B. Zoph and Q. V . Le, “Neural architecture search with reinforce- ment learning,” in Proc. Int. Conf. Learn. Repren., 2017. 2
2017
-
[53]
Efficient neural architecture search via parameter sharing,
H. Pham, M. Y. Guan, B. Zoph, Q. V . Le, and J. Dean, “Efficient neural architecture search via parameter sharing,” in Proc. Int. Conf. Mach. Learn., 2018. 2
2018
-
[54]
Learning trans- ferable architectures for scalable image recognition,
B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning trans- ferable architectures for scalable image recognition,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018. 2
2018
-
[55]
Progressive neural architecture search,
C. Liu, B. Zoph, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2
2018
-
[56]
Regularized evolution for image classifier architecture search,
E. Real, A. Aggarwal, Y. Huang, and Q. V . Le, “Regularized evolution for image classifier architecture search,” in Proceedings of the aaai conference on artificial intelligence , vol. 33, 2019, pp. 4780–
2019
-
[57]
Darts: Differentiable architec- ture search,
H. Liu, K. Simonyan, and Y. Yang, “Darts: Differentiable architec- ture search,” in Proc. Int. Conf. Learn. Repren., 2019. 2
2019
-
[58]
Proxylessnas: Direct neural architec- ture search on target task and hardware,
H. Cai, L. Zhu, and S. Han, “Proxylessnas: Direct neural architec- ture search on target task and hardware,” in Proc. Int. Conf. Learn. Repren., 2019. 2
2019
-
[59]
Nisp: Pruning networks using neuron importance score propagation,
R. Yu, A. Li, C.-F. Chen, J.-H. Lai, V . I. Morariu, X. Han, M. Gao, C.-Y. Lin, and L. S. Davis, “Nisp: Pruning networks using neuron importance score propagation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018, pp. 9194–9203. 2
2018
-
[60]
N2n learning: Network to network compression via policy gradient reinforcement learning,
A. Ashok, N. Rhinehart, F. Beainy, and K. M. Kitani, “N2n learning: Network to network compression via policy gradient reinforcement learning,” in Proc. Int. Conf. Learn. Repren., 2018. 2
2018
-
[61]
Amc: Automl for model compression and acceleration on mobile devices,
Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “Amc: Automl for model compression and acceleration on mobile devices,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018. 2
2018
-
[62]
Clip-q: Deep network compression learning by in-parallel pruning-quantization,
F. Tung and G. Mori, “Clip-q: Deep network compression learning by in-parallel pruning-quantization,” inProc. IEEE Conf. Comp. Vis. Patt. Recogn., 2018, pp. 7873–7882. 2
2018
-
[63]
Network pruning via transformable archi- tecture search,
X. Dong and Y. Yang, “Network pruning via transformable archi- tecture search,” in Proc. Adv. Neural Inf. Process. Syst. , 2019, pp. 760–771. 2
2019
-
[64]
Real-time action recognition with enhanced motion vector cnns,
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang, “Real-time action recognition with enhanced motion vector cnns,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2016, pp. 2718–2726. 3
2016
-
[65]
Learning efficient object detection models with knowledge distillation,
G. Chen, W. Choi, X. Yu, T. Han, and M. Chandraker, “Learning efficient object detection models with knowledge distillation,” in Proc. Adv. Neural Inf. Process. Syst., 2017, pp. 742–751. 3
2017
-
[66]
Quantization mimic: Towards very tiny cnn for object detection,
Y. Wei, X. Pan, H. Qin, W. Ouyang, and J. Yan, “Quantization mimic: Towards very tiny cnn for object detection,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2018, pp. 267–283. 3
2018
-
[67]
Knowledge adaptation for efficient semantic segmentation,
T. He, C. Shen, Z. Tian, D. Gong, C. Sun, and Y. Yan, “Knowledge adaptation for efficient semantic segmentation,” arXiv preprint arXiv:1903.04688, 2019. 3
1903 arXiv
-
[68]
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,
A. Mishra and D. Marr, “Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,” in Proc. Int. Conf. Learn. Repren., 2018. 3
2018
-
[69]
Maxout networks,
I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, “Maxout networks,” in Proc. Int. Conf. Mach. Learn. ,
-
[70]
Regular- ization of neural networks using dropconnect,
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus, “Regular- ization of neural networks using dropconnect,” in Proc. Int. Conf. Mach. Learn., 2013, pp. 1058–1066. 3
2013
-
[71]
Gradual dropin of layers to train very deep neural networks,
L. N. Smith, E. M. Hand, and T. Doster, “Gradual dropin of layers to train very deep neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2016, pp. 4763–4771. 3
2016
-
[72]
Learning accurate low-bit deep neural networks with stochastic quantization,
Y. Dong, R. Ni, J. Li, Y. Chen, J. Zhu, and H. Su, “Learning accurate low-bit deep neural networks with stochastic quantization,” in Proc. Brit. Mach. Vis. Conf., 2017. 3
2017
-
[73]
Slimmable neural networks,
J. Yu, L. Yang, N. Xu, J. Yang, and T. Huang, “Slimmable neural networks,” 2019. 3, 4
2019
-
[74]
Universally slimmable networks and improved training techniques,
J. Yu and T. S. Huang, “Universally slimmable networks and improved training techniques,” in Proc. IEEE Int. Conf. Comp. Vis., 2019, pp. 1803–1811. 3
2019
-
[75]
Learning multiple layers of fea- tures from tiny images,
A. Krizhevsky and G. Hinton, “Learning multiple layers of fea- tures from tiny images,” 2009. 6
2009
-
[76]
Imagenet large scale visual recognition challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” Int. J. Comp. Vis., vol. 115, no. 3, pp. 211–252, 2015. 6
2015
-
[77]
Identity mappings in deep residual networks,
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in Proc. Eur. Conf. Comp. Vis. Workshops, 2016, pp. 630–645. 6
2016
-
[78]
Precision highway for ultra low-precision quantization,
E. Park, D. Kim, S. Yoo, and P . Vajda, “Precision highway for ultra low-precision quantization,” arXiv preprint arXiv:1812.09818, 2018. 8 Bohan Zhuang is a Lecturer at Monash University, Australia. He re- ceived his PhD degree from The University of Adelaide. Mingkui Tan is a...
2018 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.