REVIEW 4 cited by
LongSafety: Enhance Safety for Long-Context LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for their application in increasingly complex tasks. However, despite the growing capabilities of long-context LLMs, the safety issues in long-context scenarios remain underexplored. While safety alignment in short context has been widely studied, the safety concerns of long-context LLMs have not been adequately addressed. In this work, we introduce \textbf{LongSafety}, a comprehensive safety alignment dataset for long-context LLMs, containing 10 tasks and 17k samples, with an average length of 40.9k tokens. Our experiments demonstrate that training with LongSafety can enhance long-context safety performance while enhancing short-context safety and preserving general capabilities. Furthermore, we demonstrate that long-context safety does not equal long-context alignment with short-context safety data and LongSafety has generalizing capabilities in context length and long-context safety scenarios.
Forward citations
Cited by 4 Pith papers
-
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models
Zero-RL multi-stage constructive safety alignment with SERL and long-context training lets a 14B model match much larger models on safety without collapsing helpfulness or style.
-
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
Risk-adaptive activation steering, guided by a prototype-similarity risk score computed on the first three response tokens, substantially reduces multimodal jailbreak success rates across four MLLMs while preserving utility.
-
o3-mini vs DeepSeek-R1: Which One is Safer?
DeepSeek-R1 (70B) produced unsafe responses to 11.98% of 1,260 unsafe test prompts, while OpenAI's o3-mini beta produced 1.19%, though the comparison is system-level due to API guardrails.
-
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
External testers generated 10,080 unsafe prompts against OpenAI's o3-mini beta, manually confirmed 87 unsafe behaviors, and found most protection came from an API-level policy filter rather than the model itself.
Discussion (0). Continue with ORCID to comment.