REVIEW 4 major objections 6 minor 47 references
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation
T0 review · 4 major / 6 minor · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Calibrating which feature correlations to suppress lets small models learn better from teachers with different architectures.
desk verdict Solid incremental fix on RSD: calibrated off-diagonal decorrelation plus loss-scale balancing, with clear CIFAR gains and thinner ImageNet support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Semantic Correlation Calibration (SCC): the off-diagonal part of the teacher–student correlation objective is reweighted by confusion-derived sample strengths (CEE) and a classifier-subspace strength map (SAC), with Adaptive Coefficient Regulation (ACR) setting the SCC weight from the running SCC-to-task loss ratio.
What would settle it
If ablating CEE and SAC (reverting to uniform off-diagonal penalties) or replacing the classifier-derived strength map and retrieved confusion bank with random counterparts closes the accuracy gap to RSD on the same heterogeneous CIFAR-100 and ImageNet pairs, the calibration claim fails.
Extended reading notes
Core claim
Uniform off-diagonal decorrelation in teacher–student feature correlations is too blunt for heterogeneous knowledge distillation: some correlations encode structural information that should be protected. CoCaRS improves distillation by calibrating that decorrelation—via sample-level confusion weights from teacher dark knowledge and a semantic strength map from the teacher classifier subspace—while adaptively regulating the term’s contribution from its loss scale relative to cross-entropy, yielding higher student accuracy and lower coefficient sensitivity than fixed uniform redundancy suppression.
Load-bearing premise
The method assumes that teacher-response confusion scores and associations in the teacher classifier’s feature subspace correctly mark which correlations are structure to keep rather than redundancy to kill.
Editorial extensions
If this is right
- Heterogeneous distillation should treat off-diagonal feature correlations as mixed structure-plus-redundancy, not pure noise to zero uniformly.
- Teacher classifier geometry and non-target logits can serve as practical guides for how hard to decorrelate each feature pair.
- Loss-scale adaptive weighting can replace brittle fixed coefficients when a distillation auxiliary term changes magnitude across architectures and epochs.
- Students distilled this way should show higher intermediate-feature similarity to heterogeneous teachers, especially where uniform RSD still leaves a gap.
Reading between the lines
- The same protect-structure-while-decorrelating idea may transfer to other cross-architecture alignment settings beyond classification KD, such as detection or multi-modal towers with mismatched backbones.
- If classifier-subspace maps are the right structural prior mainly near neural-collapse regimes, calibration quality may track how collapsed the teacher is, suggesting a simple diagnostic before distillation.
- Building a retrieval bank of teacher logits adds a preprocessing step; cheaper online approximations of confusion evidence would test how much of the gain needs the offline bank.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CoCaRS for heterogeneous knowledge distillation, refining RSD-style redundancy suppression. It keeps diagonal cross-architecture invariance in the teacher–student Pearson correlation matrix while calibrating off-diagonal decorrelation via Semantic Correlation Calibration (SCC): Confusion Evidence Estimation (CEE) builds sample confusion weights from positive/reciprocal-negative teacher dark knowledge in a retrieval bank (Eqs. 5–8), and Strength Allocation Control (SAC) builds a semantic strength map from a QR-induced discriminative subspace of normalized teacher classifier weights so stronger associations receive weaker decorrelation (Eq. 9). Adaptive Coefficient Regulation (ACR) then sets the SCC coefficient from the relative SCC/CE loss scale via EMA (Eqs. 13–15). On CIFAR-100 and ImageNet-1K across CNN/ViT/MLP pairs, CoCaRS reports higher Top-1 than RSD and other heterogeneous KD baselines (CIFAR avg. 84.70% vs RSD 82.14%; ImageNet avg. 74.46% vs 74.03%), with component and formulation ablations, homogeneous transfer, CKA similarity, and coefficient-scale plots.
Significance. If the gains hold under stronger statistical reporting and clearer mechanism checks, CoCaRS is a useful incremental advance on redundancy-suppression heterogeneous KD: it targets two practical RSD weaknesses (uniform off-diagonal decorrelation and fixed β) with a modular calibration stack and broad multi-architecture evaluation. Strengths include systematic ablations (Tables 3–6), homogeneous-setting results (Table 7), CKA analysis (Fig. 2), compute/memory comparison (Fig. 3), and explicit ACR dynamics (Fig. 4). The work is empirical rather than theoretical; significance rests on reliable large-scale gains and on whether CEE/SAC truly separate structure from redundancy rather than acting as useful but opaque regularizers. Code release is promised but not yet available.
major comments (4)
- [Method, Strength Allocation Control; Table 5] SAC’s signed structure-vs-redundancy assumption is load-bearing but under-justified. Method (SAC, Eq. 9) sets M_κ ∝ exp(−τ_κ M_off) with M_sem = |QQᵀ| from QR on normalized teacher classifier weights, so larger off-diagonal subspace association lowers decorrelation. Neural-collapse motivation is ambiguous: co-loading dimensions may be redundant copies of one discriminative direction (favoring stronger decorrelation) rather than complementary structure. Table 5 shows Inverted ≈ w/o SAC and Random worse than full SAC, which establishes usefulness of the knob and preferred sign under this recipe, not that M_off identifies structural correlations that should be preserved. A direct diagnostic is needed (e.g., controlled synthetic redundancy, or measuring whether protected correlations improve class-separability / NC geometry rather than only final accuracy).
- [Method, Confusion Evidence Estimation; Table 4] CEE may confound “calibrated RSD” with auxiliary multi-teacher dark-knowledge aggregation. Eqs. 4–8 build a response bank over pretrained teachers {T_m}, retrieve neighbors with a key encoder, and reweight correlation estimation by w_conf. Gains attributed to correlation calibration could partly come from this external supervision channel rather than better structure/redundancy separation inside P. Table 4’s Random retrieval drop helps, but does not isolate bank/multi-teacher logits from pure sample reweighting of the existing teacher–student pair. Please ablate single-teacher bank vs multi-teacher, bank-free CEE (e.g., using only the current teacher’s non-target logits), and report whether SCC still beats RSD without retrieval infrastructure.
- [Table 2; Experiments] ImageNet support for the central claim is thin relative to the CIFAR story. Table 2 average lift over RSD is +0.43% (one pair −0. something on Mixer→MobileNetV2 where CoCaRS 72.15 is below PAT 72.22), with no error bars, seeds, or significance tests anywhere in Tables 1–2. The strongest claim couples CIFAR and ImageNet improvements plus reduced coefficient sensitivity; without multi-seed uncertainty, the large-scale half is not yet commensurate with the mechanism narrative. At minimum report mean±std over ≥3 seeds on ImageNet pairs and clarify hyperparameter selection for κ, ρ, α, γ, τ_κ.
- [Abstract; Adaptive Coefficient Regulation; Fig. 4] ACR’s claim to “reduce sensitivity to coefficient settings” is only partially evidenced. Fig. 4 shows relative SCC/CE scale varies across pairs and that λ_t is adapted, but there is no head-to-head sensitivity sweep (fixed λ grid vs ACR) reporting accuracy variance across coefficients/pairs/stages—the quantity the abstract and introduction emphasize. Without that comparison (or a table of performance under mismatched fixed β/λ), the sensitivity-reduction claim remains qualitative. Please add a coefficient-sensitivity experiment on at least two heterogeneous pairs.
minor comments (6)
- [Preliminaries, Eq. (1)] Eq. (1) formatting is hard to parse (missing clear fraction bars/parentheses for the Pearson denominator). Please re-typeset P_ij in standard form.
- [Fig. 4] Fig. 4 caption/text inconsistency: body says “After ARC” in one place while the method is ACR; fix typo.
- [Method; Experiments] Several free parameters (κ, ρ, η, G(·), α, γ, τ_κ, retrieval design, stable-rank target) are introduced with limited default values or selection protocol in the main text. A compact hyperparameter table would aid reproducibility pending code release.
- [Related Work] Related Work could more sharply separate CoCaRS from SemCKD / SimKD / NCKD beyond motivation, since CEE and classifier-subspace use overlap thematically with those lines.
- [Table 1] Table 1 DIST entry for Swin-T→ResMLP-S12 (11.05%) looks like a training failure; flag or explain outliers so averages are not skewed without comment.
- [Abstract] Promise that “Code will be released soon” should be paired with a concrete artifact plan (configs, bank construction, seeds) given the retrieval-bank dependency.
Circularity Check
No significant circularity: empirical KD method with external benchmarks; losses and proxies are design choices, not self-fulfilling predictions.
full rationale
CoCaRS is an empirical methods paper. The claimed gains (e.g., CIFAR-100 avg 84.70% vs RSD 82.14%; ImageNet-1K 74.46% vs 74.03%) are measured Top-1 accuracies against external baselines and ablations, not quantities algebraically forced by the loss definitions. SCC (CEE Eqs. 5–8, SAC Eq. 9, L_SCC Eq. 10), ACR (r_t, λ_t Eqs. 13–14), and the overall objective L = L_CE + λ_t L_SCC are proposed regularizers and coefficient schedules; training then optimizes the student, and reported accuracy is an independent evaluation outcome. Neural-collapse / classifier-subspace motivation and the RSD baseline are cited as external prior work, not as author uniqueness theorems that forbid alternatives or smuggle the result. ACR’s target ratio ρ is an explicit design hyperparameter, not a fitted input renamed as a prediction. No step reduces a claimed first-principles result to its own inputs by construction. Concerns about whether M_sem and confusion weights correctly separate structure from redundancy are assumption/correctness risks, not circularity.
Assumptions & free parameters
free parameters (4)
- κ (overall decorrelation strength in L_SCC)
- ACR target ratio ρ and EMA decay η / modulation G(·)
- CEE scalars α, γ and retrieval bank design
- SAC temperature τ_κ and stable-rank target for QR
assumptions (5)
- domain assumption Pearson correlations between adapted student features and teacher features are a valid interface for cross-architecture knowledge transfer.
- domain assumption Driving diagonal correlations to 1 preserves useful cross-architecture invariance while off-diagonal mass is largely redundancy unless semantically protected.
- ad hoc to paper Non-target teacher logits in a retrieved local neighborhood encode reliable semantic confusion structure for sample weighting.
- ad hoc to paper Associations in the teacher classifier’s discriminative subspace (via QR on weights) mark feature-dimension pairs whose correlations should be decorrelated less.
- domain assumption Standard supervised vision benchmarks (CIFAR-100, ImageNet-1K Top-1) and the listed baselines are adequate to judge heterogeneous KD methods.
invented entities (3)
-
Confusion evidence weights w_conf from asymmetric positive/reciprocal-negative teacher evidence
-
Semantic strength map M_κ from teacher-classifier QR subspace
-
Adaptive Coefficient Regulation (ACR) state λ_t
Cite this review
Pith. "Pith review of CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation." pith.science (2026). https://pith.science/paper/RNDHDAEM
@misc{pith2026260727054,
author = {Pith},
title = {Pith review of: CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/RNDHDAEM}},
note = {Machine review of arXiv:2607.27054}
}
read the original abstract
Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stages. Specifically, CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC), which respectively capture reliable semantic relations for correlation estimation and preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of the calibrated redundancy suppression objective according to its relative loss scale, reducing sensitivity to coefficient settings. Extensive experiments on CIFAR-100 and ImageNet-1K validate the effectiveness of CoCaRS in improving distillation performance and reducing sensitivity to coefficient settings. Code will be released soon.
Figures
Reference graph
Works this paper leans on
-
[1]
Bardes, A.; Ponce, J.; and LeCun, Y. 2022. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. In The Tenth International Conference on Learning Representations, ICLR 2022 . OpenReview.net
2022
-
[2]
Chen, D.; Mei, J.; Zhang, H.; Wang, C.; Feng, Y.; and Chen, C. 2022. Knowledge Distillation with the Reused Teacher Classifier. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11923--11932. IEEE
2022
-
[3]
Chen, D.; Mei, J.; Zhang, Y.; Wang, C.; Wang, Z.; Feng, Y.; and Chen, C. 2021 a . Cross-Layer Distillation with Semantic Calibration. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021 , 7028--7036. AAAI Press
2021
-
[4]
Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021 b . Distilling Knowledge via Knowledge Review. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021 , 5008--5017. Computer Vision Foundation / IEEE
2021
-
[5]
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248--255. Ieee
2009
-
[6]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations, ICLR 2021 . OpenReview.net
2021
-
[7]
Gou, J.; Yu, B.; and Maybank, S. J. 2021. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6): 1789--1819
2021
-
[8]
Guo, Z.; Yan, H.; Li, H.; and Lin, X. 2023. Class Attention Transfer Based Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 11868--11877. IEEE
2023
Show all 47 references
-
[9]
Hao, Z.; Guo, J.; Han, K.; Tang, Y.; Hu, H.; Wang, Y.; and Xu, C. 2023. One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing System...
2023
-
[10]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016 , 770--778. IEEE Computer Society
2016
-
[11]
Heo, B.; Kim, J.; Yun, S.; Park, H.; Kwak, N.; and Choi, J. Y. 2019 a . A Comprehensive Overhaul of Feature Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 1921--1930. IEEE
2019
-
[12]
Heo, B.; Lee, M.; Yun, S.; and Choi, J. Y. 2019 b . Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019 , 3779--3787. AAAI Press
2019
-
[13]
E.; Vinyals, O.; and Dean, J
Hinton, G. E.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. CoRR, abs/1503.02531
2015 arXiv
-
[14]
Huang, T.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2022. Knowledge Distillation from A Stronger Teacher. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022
2022
-
[15]
Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019. Similarity of Neural Network Representations Revisited. In Proceedings of the 36th International Conference on Machine Learning, volume 97, 3519--3529. PMLR
2019
-
[16]
Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images
2009
-
[17]
Lao, S.; Song, G.; Liu, B.; Liu, Y.; and Yang, Y. 2023. Masked Autoencoders Are Stronger Knowledge Distillers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , 6361--6370. IEEE
2023
-
[18]
Li, G.; Wang, Q.; Yan, K.; Ding, S.; Gao, Y.; and Xia, G. 2025. Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025 , 3445--3454. IEEE
2025
-
[19]
Lin, J.; Yao, Y.; Hsu, C.; Xie, H.; Shuai, H.; and Cheng, W. 2025. Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025 , 4178--4187. IEEE
2025
-
[20]
Lin, S.; Xie, H.; Wang, B.; Yu, K.; Chang, X.; Liang, X.; and Wang, G. 2022. Knowledge Distillation via the Target-aware Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 10905--10914. IEEE
2022
-
[21]
Liu, L.; Huang, Q.; Lin, S.; Xie, H.; Wang, B.; Chang, X.; and Liang, X. 2021 a . Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 8251--8260. IEEE
2021
-
[22]
Liu, Y.; Cao, J.; Li, B.; Hu, W.; Ding, J.; and Li, L. 2022 a . Cross-Architecture Knowledge Distillation. In Computer Vision - ACCV 2022 , volume 13845, 179--195. Springer
2022
-
[23]
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 b . Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 9992--10002. IEEE
2021
-
[24]
Liu, Z.; Mao, H.; Wu, C.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022 b . A ConvNet for the 2020s. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11966--11976. IEEE
2022
-
[25]
Pan, H.; Yu, F.; Zhang, K.; Lan, H.; Meng, Q.; and Li, Z. 2026. Knowledge Distillation in Visual Algorithms: A Survey. Journal of Computer Research and Development, 63(1): 90--122
2026
-
[26]
Y.; and Donoho, D
Papyan, V.; Han, X. Y.; and Donoho, D. L. 2020. Prevalence of Neural Collapse during the terminal phase of deep learning training. CoRR, abs/2008.08186
2020 arXiv
-
[27]
Park, W.; Kim, D.; Lu, Y.; and Cho, M. 2019. Relational Knowledge Distillation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019 , 3967--3976. Computer Vision Foundation / IEEE
2019
-
[28]
Peng, B.; Jin, X.; Li, D.; Zhou, S.; Wu, Y.; Liu, J.; Zhang, Z.; and Liu, Y. 2019. Correlation Congruence for Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 5006--5015. IEEE
2019
-
[29]
Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021. Do Vision Transformers See Like Convolutional Neural Networks? In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021...
2021
-
[30]
Ren, S.; Gao, Z.; Hua, T.; Xue, Z.; Tian, Y.; He, S.; and Zhao, H. 2022. Co-advise: Cross Inductive Bias Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 16752--16761. IEEE
2022
-
[31]
E.; Chassang, A.; Gatta, C.; and Bengio, Y
Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2015. FitNets: Hints for Thin Deep Nets. In 3rd International Conference on Learning Representations, ICLR 2015
2015
-
[32]
G.; Zhu, M.; Zhmoginov, A.; and Chen, L
Sandler, M.; Howard, A. G.; Zhu, M.; Zhmoginov, A.; and Chen, L. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , 4510--4520. Computer Vision Foundation / IEEE Computer Society
2018
-
[33]
Son, W.; Na, J.; Choi, J.; and Hwang, W. 2021. Densely Guided Knowledge Distillation using Multiple Teacher Assistants. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 9375--9384. IEEE
2021
-
[34]
Tian, Y.; Krishnan, D.; and Isola, P. 2020. Contrastive Representation Distillation. In 8th International Conference on Learning Representations, ICLR 2020 . OpenReview.net
2020
-
[35]
O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; Lucic, M.; and Dosovitskiy, A
Tolstikhin, I. O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; Lucic, M.; and Dosovitskiy, A. 2021. MLP-Mixer: An all-MLP Architecture for Vision. In Advances in Neural Information Processing Systems 34:...
2021
-
[36]
Touvron, H.; Bojanowski, P.; Caron, M.; Cord, M.; El - Nouby, A.; Grave, E.; Izacard, G.; Joulin, A.; Synnaeve, G.; Verbeek, J.; and J \' e gou, H. 2023. ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training. IEEE Trans. Pattern Anal. Mach. Intell....
2023
-
[37]
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and J \' e gou, H. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021 , volume 139, 10347--10357. PMLR
2021
-
[38]
Tung, F.; and Mori, G. 2019. Similarity-Preserving Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 1365--1374. IEEE
2019
-
[39]
Wei, S.; Luo, C.; and Luo, Y. 2024. Scale Decoupled Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 15975--15983. IEEE
2024
-
[40]
Xu, H.; Fang, J.; Zhang, X.; Xie, L.; Wang, X.; Dai, W.; Xiong, H.; and Tian, Q. 2022. Bag of Instances Aggregation Boosts Self-supervised Distillation. In The Tenth International Conference on Learning Representations, ICLR 2022 . OpenReview.net
2022
-
[41]
Yang, C.; Xie, L.; Qiao, S.; and Yuille, A. L. 2019. Training Deep Neural Networks in Generations: A More Tolerant Teacher Educates Better Students. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019 , 5628--5635. AAAI Press
2019
-
[42]
Yim, J.; Joo, D.; Bae, J.; and Kim, J. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 , 7130--7138. IEEE Computer Society
2017
-
[43]
Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; and Deny, S. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021 , volume 139, 12310--12320. PMLR
2021
-
[44]
Zhang, S.; Song, Z.; and He, K. 2025. Neural Collapse Inspired Knowledge Distillation. In Thirty-Ninth AAAI Conference on Artificial Intelligence, AAAI 2025 , 22542--22550. AAAI Press
2025
-
[45]
Zhang, W.; Liu, Y.; Ran, W.; and Ma, C. 2025. Cross-Architecture Distillation Made Simple with Redundancy Suppression. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, October 19-25, 2025 , 23256--23266. IEEE
2025
-
[46]
Zhao, B.; Cui, Q.; Song, R.; Qiu, Y.; and Liang, J. 2022. Decoupled Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11943--11952. IEEE
2022
-
[47]
Zhao, B.; Song, R.; and Liang, J. 2023. Cumulative Spatial Knowledge Distillation for Vision Transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , 6123--6132. IEEE
2023
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.