Pith. sign in

REVIEW 4 major objections 6 minor 47 references

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

T0 review · 4 major / 6 minor · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Calibrating which feature correlations to suppress lets small models learn better from teachers with different architectures.

desk verdict Solid incremental fix on RSD: calibrated off-diagonal decorrelation plus loss-scale balancing, with clear CIFAR gains and thinner ImageNet support. read the letter →

arxiv 2607.27054 v1 pith:RNDHDAEM submitted 2026-07-29 cs.LG cs.AI

classification cs.LGcs.AI
keywords knowledgedistillationheterogeneousarchitecturesredundancysuppressionfeaturecorrelationcalibrationmodelcompressionadaptivelossweighting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

When a compact student network is trained from a larger teacher with a different architecture, their internal features often disagree in ways that blunt ordinary knowledge distillation. Prior redundancy-suppression methods force teacher–student feature correlations toward identity, keeping cross-architecture invariance on the diagonal while zeroing off-diagonal correlations as pure redundancy. This paper argues that uniform zeroing also erases useful structure, and that a fixed loss weight makes the method brittle across model pairs and training stages. CoCaRS instead calibrates decorrelation: confusion evidence from teacher responses reweights which samples matter for estimating correlations, a strength map built from the teacher classifier’s discriminative subspace softens decorrelation where structure should be kept, and an adaptive coefficient keeps the calibrated term balanced against the task loss. On CIFAR-100 and ImageNet-1K across CNN, ViT, and MLP pairs, the approach raises accuracy over uniform redundancy suppression and prior heterogeneous distillation baselines while reducing sensitivity to the coefficient.

What carries the argument

Semantic Correlation Calibration (SCC): the off-diagonal part of the teacher–student correlation objective is reweighted by confusion-derived sample strengths (CEE) and a classifier-subspace strength map (SAC), with Adaptive Coefficient Regulation (ACR) setting the SCC weight from the running SCC-to-task loss ratio.

What would settle it

If ablating CEE and SAC (reverting to uniform off-diagonal penalties) or replacing the classifier-derived strength map and retrieved confusion bank with random counterparts closes the accuracy gap to RSD on the same heterogeneous CIFAR-100 and ImageNet pairs, the calibration claim fails.

Watch

Extended reading notes

Core claim

Uniform off-diagonal decorrelation in teacher–student feature correlations is too blunt for heterogeneous knowledge distillation: some correlations encode structural information that should be protected. CoCaRS improves distillation by calibrating that decorrelation—via sample-level confusion weights from teacher dark knowledge and a semantic strength map from the teacher classifier subspace—while adaptively regulating the term’s contribution from its loss scale relative to cross-entropy, yielding higher student accuracy and lower coefficient sensitivity than fixed uniform redundancy suppression.

Load-bearing premise

The method assumes that teacher-response confusion scores and associations in the teacher classifier’s feature subspace correctly mark which correlations are structure to keep rather than redundancy to kill.

Editorial extensions

If this is right

  • Heterogeneous distillation should treat off-diagonal feature correlations as mixed structure-plus-redundancy, not pure noise to zero uniformly.
  • Teacher classifier geometry and non-target logits can serve as practical guides for how hard to decorrelate each feature pair.
  • Loss-scale adaptive weighting can replace brittle fixed coefficients when a distillation auxiliary term changes magnitude across architectures and epochs.
  • Students distilled this way should show higher intermediate-feature similarity to heterogeneous teachers, especially where uniform RSD still leaves a gap.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same protect-structure-while-decorrelating idea may transfer to other cross-architecture alignment settings beyond classification KD, such as detection or multi-modal towers with mismatched backbones.
  • If classifier-subspace maps are the right structural prior mainly near neural-collapse regimes, calibration quality may track how collapsed the teacher is, suggesting a simple diagnostic before distillation.
  • Building a retrieval bank of teacher logits adds a preprocessing step; cheaper online approximations of confusion evidence would test how much of the gain needs the offline bank.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes CoCaRS for heterogeneous knowledge distillation, refining RSD-style redundancy suppression. It keeps diagonal cross-architecture invariance in the teacher–student Pearson correlation matrix while calibrating off-diagonal decorrelation via Semantic Correlation Calibration (SCC): Confusion Evidence Estimation (CEE) builds sample confusion weights from positive/reciprocal-negative teacher dark knowledge in a retrieval bank (Eqs. 5–8), and Strength Allocation Control (SAC) builds a semantic strength map from a QR-induced discriminative subspace of normalized teacher classifier weights so stronger associations receive weaker decorrelation (Eq. 9). Adaptive Coefficient Regulation (ACR) then sets the SCC coefficient from the relative SCC/CE loss scale via EMA (Eqs. 13–15). On CIFAR-100 and ImageNet-1K across CNN/ViT/MLP pairs, CoCaRS reports higher Top-1 than RSD and other heterogeneous KD baselines (CIFAR avg. 84.70% vs RSD 82.14%; ImageNet avg. 74.46% vs 74.03%), with component and formulation ablations, homogeneous transfer, CKA similarity, and coefficient-scale plots.

Significance. If the gains hold under stronger statistical reporting and clearer mechanism checks, CoCaRS is a useful incremental advance on redundancy-suppression heterogeneous KD: it targets two practical RSD weaknesses (uniform off-diagonal decorrelation and fixed β) with a modular calibration stack and broad multi-architecture evaluation. Strengths include systematic ablations (Tables 3–6), homogeneous-setting results (Table 7), CKA analysis (Fig. 2), compute/memory comparison (Fig. 3), and explicit ACR dynamics (Fig. 4). The work is empirical rather than theoretical; significance rests on reliable large-scale gains and on whether CEE/SAC truly separate structure from redundancy rather than acting as useful but opaque regularizers. Code release is promised but not yet available.

major comments (4)
  1. [Method, Strength Allocation Control; Table 5] SAC’s signed structure-vs-redundancy assumption is load-bearing but under-justified. Method (SAC, Eq. 9) sets M_κ ∝ exp(−τ_κ M_off) with M_sem = |QQᵀ| from QR on normalized teacher classifier weights, so larger off-diagonal subspace association lowers decorrelation. Neural-collapse motivation is ambiguous: co-loading dimensions may be redundant copies of one discriminative direction (favoring stronger decorrelation) rather than complementary structure. Table 5 shows Inverted ≈ w/o SAC and Random worse than full SAC, which establishes usefulness of the knob and preferred sign under this recipe, not that M_off identifies structural correlations that should be preserved. A direct diagnostic is needed (e.g., controlled synthetic redundancy, or measuring whether protected correlations improve class-separability / NC geometry rather than only final accuracy).
  2. [Method, Confusion Evidence Estimation; Table 4] CEE may confound “calibrated RSD” with auxiliary multi-teacher dark-knowledge aggregation. Eqs. 4–8 build a response bank over pretrained teachers {T_m}, retrieve neighbors with a key encoder, and reweight correlation estimation by w_conf. Gains attributed to correlation calibration could partly come from this external supervision channel rather than better structure/redundancy separation inside P. Table 4’s Random retrieval drop helps, but does not isolate bank/multi-teacher logits from pure sample reweighting of the existing teacher–student pair. Please ablate single-teacher bank vs multi-teacher, bank-free CEE (e.g., using only the current teacher’s non-target logits), and report whether SCC still beats RSD without retrieval infrastructure.
  3. [Table 2; Experiments] ImageNet support for the central claim is thin relative to the CIFAR story. Table 2 average lift over RSD is +0.43% (one pair −0. something on Mixer→MobileNetV2 where CoCaRS 72.15 is below PAT 72.22), with no error bars, seeds, or significance tests anywhere in Tables 1–2. The strongest claim couples CIFAR and ImageNet improvements plus reduced coefficient sensitivity; without multi-seed uncertainty, the large-scale half is not yet commensurate with the mechanism narrative. At minimum report mean±std over ≥3 seeds on ImageNet pairs and clarify hyperparameter selection for κ, ρ, α, γ, τ_κ.
  4. [Abstract; Adaptive Coefficient Regulation; Fig. 4] ACR’s claim to “reduce sensitivity to coefficient settings” is only partially evidenced. Fig. 4 shows relative SCC/CE scale varies across pairs and that λ_t is adapted, but there is no head-to-head sensitivity sweep (fixed λ grid vs ACR) reporting accuracy variance across coefficients/pairs/stages—the quantity the abstract and introduction emphasize. Without that comparison (or a table of performance under mismatched fixed β/λ), the sensitivity-reduction claim remains qualitative. Please add a coefficient-sensitivity experiment on at least two heterogeneous pairs.
minor comments (6)
  1. [Preliminaries, Eq. (1)] Eq. (1) formatting is hard to parse (missing clear fraction bars/parentheses for the Pearson denominator). Please re-typeset P_ij in standard form.
  2. [Fig. 4] Fig. 4 caption/text inconsistency: body says “After ARC” in one place while the method is ACR; fix typo.
  3. [Method; Experiments] Several free parameters (κ, ρ, η, G(·), α, γ, τ_κ, retrieval design, stable-rank target) are introduced with limited default values or selection protocol in the main text. A compact hyperparameter table would aid reproducibility pending code release.
  4. [Related Work] Related Work could more sharply separate CoCaRS from SemCKD / SimKD / NCKD beyond motivation, since CEE and classifier-subspace use overlap thematically with those lines.
  5. [Table 1] Table 1 DIST entry for Swin-T→ResMLP-S12 (11.05%) looks like a training failure; flag or explain outliers so averages are not skewed without comment.
  6. [Abstract] Promise that “Code will be released soon” should be paired with a concrete artifact plan (configs, bank construction, seeds) given the retrieval-bank dependency.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical KD method with external benchmarks; losses and proxies are design choices, not self-fulfilling predictions.

full rationale

CoCaRS is an empirical methods paper. The claimed gains (e.g., CIFAR-100 avg 84.70% vs RSD 82.14%; ImageNet-1K 74.46% vs 74.03%) are measured Top-1 accuracies against external baselines and ablations, not quantities algebraically forced by the loss definitions. SCC (CEE Eqs. 5–8, SAC Eq. 9, L_SCC Eq. 10), ACR (r_t, λ_t Eqs. 13–14), and the overall objective L = L_CE + λ_t L_SCC are proposed regularizers and coefficient schedules; training then optimizes the student, and reported accuracy is an independent evaluation outcome. Neural-collapse / classifier-subspace motivation and the RSD baseline are cited as external prior work, not as author uniqueness theorems that forbid alternatives or smuggle the result. ACR’s target ratio ρ is an explicit design hyperparameter, not a fitted input renamed as a prediction. No step reduces a claimed first-principles result to its own inputs by construction. Concerns about whether M_sem and confusion weights correctly separate structure from redundancy are assumption/correctness risks, not circularity.

Assumptions & free parameters 4 free parameters · 5 assumptions · 3 invented entities

The central claim rests on standard KD training practice plus several modeling choices imported from RSD and representation learning: Pearson teacher–student correlations as the right transfer interface; diagonal invariance + off-diagonal decorrelation as the right inductive bias; teacher logits/neighborhood retrieval as semantic confusion evidence; teacher classifier geometry (neural-collapse-motivated subspace) as a map of which correlations are structural; and relative loss-scale EMA control as sufficient to stabilize multi-term optimization. Free knobs (κ, λ/ρ path, α, γ, τ_κ, retrieval design) are chosen rather than derived.

free parameters (4)
  • κ (overall decorrelation strength in L_SCC)
    Controls off-diagonal penalty magnitude in Eq. 10; inherited from RSD’s κ and not derived from data-generating assumptions.
  • ACR target ratio ρ and EMA decay η / modulation G(·)
    Define the desired L_SCC/L_CE operating point and update dynamics (Eqs. 13–14); chosen to reduce coefficient sensitivity rather than identified from first principles.
  • CEE scalars α, γ and retrieval bank design
    α weights confusion norms into sample importance S_conf; γ scales reciprocal negative evidence; bank keys, teachers {T_m}, and neighborhood sizes shape w_conf and are implementation choices.
  • SAC temperature τ_κ and stable-rank target for QR
    τ_κ sets how sharply M_off modulates decorrelation (Eq. 9); rank reduction before QR is a heuristic on classifier weight geometry.
assumptions (5)
  • domain assumption Pearson correlations between adapted student features and teacher features are a valid interface for cross-architecture knowledge transfer.
    Taken from RSD preliminaries (Eq. 1) and used as the substrate for all CoCaRS calibration.
  • domain assumption Driving diagonal correlations to 1 preserves useful cross-architecture invariance while off-diagonal mass is largely redundancy unless semantically protected.
    Core RSD inductive bias retained in L_SCC; Table 6 argues diagonal should not receive CEE/SAC modulation.
  • ad hoc to paper Non-target teacher logits in a retrieved local neighborhood encode reliable semantic confusion structure for sample weighting.
    CEE positive/negative evidence construction (Eqs. 5–8); supported by ablation Table 4 but not independently established outside this pipeline.
  • ad hoc to paper Associations in the teacher classifier’s discriminative subspace (via QR on weights) mark feature-dimension pairs whose correlations should be decorrelated less.
    SAC semantic matrix M_sem = |QQᵀ| motivated by neural collapse citations; direction of allocation validated only by Table 5 inverted/random controls.
  • domain assumption Standard supervised vision benchmarks (CIFAR-100, ImageNet-1K Top-1) and the listed baselines are adequate to judge heterogeneous KD methods.
    Experimental setup and main tables; conventional in the subfield.
invented entities (3)
  • Confusion evidence weights w_conf from asymmetric positive/reciprocal-negative teacher evidence
    purpose: Sample-aware reweighting inside off-diagonal correlation estimation for SCC.
    Constructed in CEE from a response bank; not a physical entity but a new algorithmic latent specific to this paper’s calibration story.
  • Semantic strength map M_κ from teacher-classifier QR subspace
    purpose: Allocate non-uniform decorrelation strength across feature-dimension pairs.
    Defined in SAC (Eq. 9); existence and usefulness are justified mainly by ablations within the same benchmarks.
  • Adaptive Coefficient Regulation (ACR) state λ_t
    purpose: Keep SCC’s effective contribution stable relative to L_CE across pairs and stages.
    Online controller (Eqs. 13–14); standard loss-balancing idea specialized to SCC, without external falsifiable prediction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation." pith.science (2026). https://pith.science/paper/RNDHDAEM

@misc{pith2026260727054,
  author       = {Pith},
  title        = {Pith review of: CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RNDHDAEM}},
  note         = {Machine review of arXiv:2607.27054}
}
read the original abstract

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stages. Specifically, CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC), which respectively capture reliable semantic relations for correlation estimation and preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of the calibrated redundancy suppression objective according to its relative loss scale, reducing sensitivity to coefficient settings. Extensive experiments on CIFAR-100 and ImageNet-1K validate the effectiveness of CoCaRS in improving distillation performance and reducing sensitivity to coefficient settings. Code will be released soon.

Figures

Figures reproduced from arXiv: 2607.27054 by the authors.

Figure 1
Figure 1. Overview of the proposed CoCaRS framework. (a) CoCaRS refines redundancy suppression for heterogeneous [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 3
Figure 3. Comparison of additional trainable parameters and [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 2
Figure 2. Intermediate representation similarity between [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 2 linked inside Pith

  1. [1]

    Bardes, A.; Ponce, J.; and LeCun, Y. 2022. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. In The Tenth International Conference on Learning Representations, ICLR 2022 . OpenReview.net

  2. [2]

    Chen, D.; Mei, J.; Zhang, H.; Wang, C.; Feng, Y.; and Chen, C. 2022. Knowledge Distillation with the Reused Teacher Classifier. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11923--11932. IEEE

  3. [3]

    Chen, D.; Mei, J.; Zhang, Y.; Wang, C.; Wang, Z.; Feng, Y.; and Chen, C. 2021 a . Cross-Layer Distillation with Semantic Calibration. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021 , 7028--7036. AAAI Press

  4. [4]

    Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021 b . Distilling Knowledge via Knowledge Review. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021 , 5008--5017. Computer Vision Foundation / IEEE

  5. [5]

    Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248--255. Ieee

  6. [6]

    Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations, ICLR 2021 . OpenReview.net

  7. [7]

    Gou, J.; Yu, B.; and Maybank, S. J. 2021. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6): 1789--1819

  8. [8]

    Guo, Z.; Yan, H.; Li, H.; and Lin, X. 2023. Class Attention Transfer Based Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 11868--11877. IEEE

Show all 47 references
  1. [9]

    Hao, Z.; Guo, J.; Han, K.; Tang, Y.; Hu, H.; Wang, Y.; and Xu, C. 2023. One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing System...

  2. [10]

    He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016 , 770--778. IEEE Computer Society

  3. [11]

    Heo, B.; Kim, J.; Yun, S.; Park, H.; Kwak, N.; and Choi, J. Y. 2019 a . A Comprehensive Overhaul of Feature Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 1921--1930. IEEE

  4. [12]

    Heo, B.; Lee, M.; Yun, S.; and Choi, J. Y. 2019 b . Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019 , 3779--3787. AAAI Press

  5. [13]

    E.; Vinyals, O.; and Dean, J

    Hinton, G. E.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. CoRR, abs/1503.02531

  6. [14]

    Huang, T.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2022. Knowledge Distillation from A Stronger Teacher. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022

  7. [15]

    Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019. Similarity of Neural Network Representations Revisited. In Proceedings of the 36th International Conference on Machine Learning, volume 97, 3519--3529. PMLR

  8. [16]

    Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images

  9. [17]

    Lao, S.; Song, G.; Liu, B.; Liu, Y.; and Yang, Y. 2023. Masked Autoencoders Are Stronger Knowledge Distillers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , 6361--6370. IEEE

  10. [18]

    Li, G.; Wang, Q.; Yan, K.; Ding, S.; Gao, Y.; and Xia, G. 2025. Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025 , 3445--3454. IEEE

  11. [19]

    Lin, J.; Yao, Y.; Hsu, C.; Xie, H.; Shuai, H.; and Cheng, W. 2025. Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025 , 4178--4187. IEEE

  12. [20]

    Lin, S.; Xie, H.; Wang, B.; Yu, K.; Chang, X.; Liang, X.; and Wang, G. 2022. Knowledge Distillation via the Target-aware Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 10905--10914. IEEE

  13. [21]

    Liu, L.; Huang, Q.; Lin, S.; Xie, H.; Wang, B.; Chang, X.; and Liang, X. 2021 a . Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 8251--8260. IEEE

  14. [22]

    Liu, Y.; Cao, J.; Li, B.; Hu, W.; Ding, J.; and Li, L. 2022 a . Cross-Architecture Knowledge Distillation. In Computer Vision - ACCV 2022 , volume 13845, 179--195. Springer

  15. [23]

    Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 b . Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 9992--10002. IEEE

  16. [24]

    Liu, Z.; Mao, H.; Wu, C.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022 b . A ConvNet for the 2020s. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11966--11976. IEEE

  17. [25]

    Pan, H.; Yu, F.; Zhang, K.; Lan, H.; Meng, Q.; and Li, Z. 2026. Knowledge Distillation in Visual Algorithms: A Survey. Journal of Computer Research and Development, 63(1): 90--122

  18. [26]

    Y.; and Donoho, D

    Papyan, V.; Han, X. Y.; and Donoho, D. L. 2020. Prevalence of Neural Collapse during the terminal phase of deep learning training. CoRR, abs/2008.08186

  19. [27]

    Park, W.; Kim, D.; Lu, Y.; and Cho, M. 2019. Relational Knowledge Distillation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019 , 3967--3976. Computer Vision Foundation / IEEE

  20. [28]

    Peng, B.; Jin, X.; Li, D.; Zhou, S.; Wu, Y.; Liu, J.; Zhang, Z.; and Liu, Y. 2019. Correlation Congruence for Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 5006--5015. IEEE

  21. [29]

    Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021. Do Vision Transformers See Like Convolutional Neural Networks? In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021...

  22. [30]

    Ren, S.; Gao, Z.; Hua, T.; Xue, Z.; Tian, Y.; He, S.; and Zhao, H. 2022. Co-advise: Cross Inductive Bias Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 16752--16761. IEEE

  23. [31]

    E.; Chassang, A.; Gatta, C.; and Bengio, Y

    Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2015. FitNets: Hints for Thin Deep Nets. In 3rd International Conference on Learning Representations, ICLR 2015

  24. [32]

    G.; Zhu, M.; Zhmoginov, A.; and Chen, L

    Sandler, M.; Howard, A. G.; Zhu, M.; Zhmoginov, A.; and Chen, L. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , 4510--4520. Computer Vision Foundation / IEEE Computer Society

  25. [33]

    Son, W.; Na, J.; Choi, J.; and Hwang, W. 2021. Densely Guided Knowledge Distillation using Multiple Teacher Assistants. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 9375--9384. IEEE

  26. [34]

    Tian, Y.; Krishnan, D.; and Isola, P. 2020. Contrastive Representation Distillation. In 8th International Conference on Learning Representations, ICLR 2020 . OpenReview.net

  27. [35]

    O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; Lucic, M.; and Dosovitskiy, A

    Tolstikhin, I. O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; Lucic, M.; and Dosovitskiy, A. 2021. MLP-Mixer: An all-MLP Architecture for Vision. In Advances in Neural Information Processing Systems 34:...

  28. [36]

    Touvron, H.; Bojanowski, P.; Caron, M.; Cord, M.; El - Nouby, A.; Grave, E.; Izacard, G.; Joulin, A.; Synnaeve, G.; Verbeek, J.; and J \' e gou, H. 2023. ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training. IEEE Trans. Pattern Anal. Mach. Intell....

  29. [37]

    Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and J \' e gou, H. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021 , volume 139, 10347--10357. PMLR

  30. [38]

    Tung, F.; and Mori, G. 2019. Similarity-Preserving Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 1365--1374. IEEE

  31. [39]

    Wei, S.; Luo, C.; and Luo, Y. 2024. Scale Decoupled Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 15975--15983. IEEE

  32. [40]

    Xu, H.; Fang, J.; Zhang, X.; Xie, L.; Wang, X.; Dai, W.; Xiong, H.; and Tian, Q. 2022. Bag of Instances Aggregation Boosts Self-supervised Distillation. In The Tenth International Conference on Learning Representations, ICLR 2022 . OpenReview.net

  33. [41]

    Yang, C.; Xie, L.; Qiao, S.; and Yuille, A. L. 2019. Training Deep Neural Networks in Generations: A More Tolerant Teacher Educates Better Students. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019 , 5628--5635. AAAI Press

  34. [42]

    Yim, J.; Joo, D.; Bae, J.; and Kim, J. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 , 7130--7138. IEEE Computer Society

  35. [43]

    Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; and Deny, S. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021 , volume 139, 12310--12320. PMLR

  36. [44]

    Zhang, S.; Song, Z.; and He, K. 2025. Neural Collapse Inspired Knowledge Distillation. In Thirty-Ninth AAAI Conference on Artificial Intelligence, AAAI 2025 , 22542--22550. AAAI Press

  37. [45]

    Zhang, W.; Liu, Y.; Ran, W.; and Ma, C. 2025. Cross-Architecture Distillation Made Simple with Redundancy Suppression. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, October 19-25, 2025 , 23256--23266. IEEE

  38. [46]

    Zhao, B.; Cui, Q.; Song, R.; Qiu, Y.; and Liang, J. 2022. Decoupled Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11943--11952. IEEE

  39. [47]

    Zhao, B.; Song, R.; and Liang, J. 2023. Cumulative Spatial Knowledge Distillation for Vision Transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , 6123--6132. IEEE

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.