Pith. sign in

REVIEW 2 cited by

How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.11668 v1 pith:XLMJAUZ7 submitted 2021-12-22 cs.CL cs.LG

classification cs.CLcs.LG
keywords pre-trainedfine-tuningadversariallanguagemodelmodelsriftanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The fine-tuning of pre-trained language models has a great success in many NLP fields. Yet, it is strikingly vulnerable to adversarial examples, e.g., word substitution attacks using only synonyms can easily fool a BERT-based sentiment analysis model. In this paper, we demonstrate that adversarial training, the prevalent defense technique, does not directly fit a conventional fine-tuning scenario, because it suffers severely from catastrophic forgetting: failing to retain the generic and robust linguistic features that have already been captured by the pre-trained model. In this light, we propose Robust Informative Fine-Tuning (RIFT), a novel adversarial fine-tuning method from an information-theoretical perspective. In particular, RIFT encourages an objective model to retain the features learned from the pre-trained model throughout the entire fine-tuning process, whereas a conventional one only uses the pre-trained weights for initialization. Experimental results show that RIFT consistently outperforms the state-of-the-arts on two popular NLP tasks: sentiment analysis and natural language inference, under different attacks across various pre-trained language models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Self-attention's local Lipschitz constant can be bounded using the attention probability distribution, and the softmax Jacobian spectral norm is shown to be at most 1/2, leading to a new robustness regularizer.

  2. Leveraging Transfer Learning to Overcome Data Limitations in Czochralski Crystal Growth

    cond-mat.mtrl-sci 2025-06

Pith tools