REVIEW 3 cited by
DPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised learning (SSL) has achieved notable success in many speech processing tasks, but the large model size and heavy computational cost hinder the deployment. Knowledge distillation trains a small student model to mimic the behavior of a large teacher model. However, the student architecture usually needs to be manually designed and will remain fixed during training, which requires prior knowledge and can lead to suboptimal performance. Inspired by recent success of task-specific structured pruning, we propose DPHuBERT, a novel task-agnostic compression method for speech SSL based on joint distillation and pruning. Experiments on SUPERB show that DPHuBERT outperforms pure distillation methods in almost all tasks. Moreover, DPHuBERT requires little training time and performs well with limited training data, making it suitable for resource-constrained applications. Our method can also be applied to various speech SSL models. Our code and models will be publicly available.
Forward citations
Cited by 3 Pith papers
-
Scaling and Distilling Transformer Models for sEMG
Vanilla transformers on the emg2qwerty dataset improve cross-user typing accuracy up to 109M parameters, and simple logit distillation recovers most of the gain in a 2.2M-parameter student.
-
Model Merging for Knowledge Editing
R-SFT plus task-vector scaling and pruning is proposed for knowledge editing, but the claimed sequential-editing advantage is not validated by the reported experiments.
-
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
Combining knowledge distillation with l0 or low-rank pruning improves compressed RNN-T ASR, and joint pruning with fine-tuning gives 8.9% and 13.4% relative WER gains over baseline.
Discussion (0). Continue with ORCID to comment.