REVIEW 4 cited by
Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Instruction fine-tuning has emerged as a critical technique for customizing Large Language Models (LLMs) to specific applications. However, recent studies have highlighted significant security vulnerabilities in fine-tuned LLMs. Existing defense efforts focus more on pre-training and post-training methods, yet there remains underexplored in in-training methods. To fill this gap, we introduce a novel secure-tuning strategy called SWAT. By analyzing how module-level parameters (e.g. Q/K/V/O) affect the security feature space drift, we identify a robust subset of modules, termed Mods_Rob. Our SWAT strategy begins by warming up Mods_Rob to capture low-level features with minimal security risks, followed by training all parameters to achieve optimal task performance. Essentially, this strategy shifts the early learning burden more from global parameters to Mods_Rob, reducing update magnitudes of the non-robust subset. Across various datasets, scenarios, and LLMs, our strategy has demonstrated significant success in mitigating security risks while preserving task performance. Importantly, it can be seamlessly integrated with pre-training and post-training methods, leading to greater improvements.
Forward citations
Cited by 4 Pith papers
-
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
TamperBench systematically benchmarks 21 open-weight LLMs against nine fine-tuning and representation-space attacks and finds all of them can be tampered into producing harmful output while retaining capability.
-
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
A representation-space reshaping method improves LALM safety against harmful audio queries while keeping over-rejection low.
-
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
CTRAP embeds a conditional failure mode during alignment so that harmful fine-tuning degrades the model to meaningless output while benign fine-tuning is unaffected.
-
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
MoGUv2 embeds small routers in the deeper layers of LLMs to dynamically blend a helpful variant and a refusal variant, improving safety against jailbreak and fine-tuning attacks while preserving usability.
Discussion (0). Continue with ORCID to comment.