Pith. sign in

REVIEW 1 cited by

Training Large ASR Encoders with Differential Privacy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13953 v1 pith:223YGI7W submitted 2024-09-21 cs.SD cs.CRcs.LGeess.AS

classification cs.SDcs.CRcs.LGeess.AS
keywords datalargeapplyextrapolationmodelspre-trainingpublicscales
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Self-supervised learning (SSL) methods for large speech models have proven to be highly effective at ASR. With the interest in public deployment of large pre-trained models, there is a rising concern for unintended memorization and leakage of sensitive data points from the training data. In this paper, we apply differentially private (DP) pre-training to a SOTA Conformer-based encoder, and study its performance on a downstream ASR task assuming the fine-tuning data is public. This paper is the first to apply DP to SSL for ASR, investigating the DP noise tolerance of the BEST-RQ pre-training method. Notably, we introduce a novel variant of model pruning called gradient-based layer freezing that provides strong improvements in privacy-utility-compute trade-offs. Our approach yields a LibriSpeech test-clean/other WER (%) of 3.78/ 8.41 with ($10$, 1e^-9)-DP for extrapolation towards low dataset scales, and 2.81/ 5.89 with (10, 7.9e^-11)-DP for extrapolation towards high scales.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities

    cs.CL 2025-02 accept novelty 4.0 of 10

    A survey and position paper mapping privacy threats in mental health AI and recommending a pipeline of anonymization, synthetic data, and differential privacy.

Pith tools