REVIEW 3 cited by
Watermarking Pre-trained Language Models with Backdooring
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large pre-trained language models (PLMs) have proven to be a crucial component of modern natural language processing systems. PLMs typically need to be fine-tuned on task-specific downstream datasets, which makes it hard to claim the ownership of PLMs and protect the developer's intellectual property due to the catastrophic forgetting phenomenon. We show that PLMs can be watermarked with a multi-task learning framework by embedding backdoors triggered by specific inputs defined by the owners, and those watermarks are hard to remove even though the watermarked PLMs are fine-tuned on multiple downstream tasks. In addition to using some rare words as triggers, we also show that the combination of common words can be used as backdoor triggers to avoid them being easily detected. Extensive experiments on multiple datasets demonstrate that the embedded watermarks can be robustly extracted with a high success rate and less influenced by the follow-up fine-tuning.
Forward citations
Cited by 3 Pith papers
-
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
SGRE extracts a reasoning skeleton from a teacher trace, coarsens its graph, and verbalizes it densely; the final answer is preserved verbatim, and students distilled on the edited traces show large accuracy drops.
-
Zero-Shot Attribution for Large Language Models: A Distribution Testing Approach
Anubis re-frames LLM attribution as a distribution testing problem with EVAL access, and reports AUROC above 0.9 on code benchmarks with around 2000 samples, beating detectGPT.
-
Invariant-based Robust Weights Watermark for Large Language Models
An invariant-based weights watermark embeds per-user keys into the null space of transformer invariants and uses noise to repel collusion.
Discussion (0). Continue with ORCID to comment.