Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:46:16.031718Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.08542.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:46:16.031718Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ea4a5626-c8c3-4e70-8829-b18cfd35a061 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs SafeMERGE head-to-head (main paper§7).We compare against the selective-layer baseline Safe- MERGE (Djuhera et al
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5b9e2e57-c5aa-42e1-ac6f-76b9d5c56953 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b2f3a2-c95f-4eb2-b658-4266ccfd58fb · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c50b2d-4fdf-4c95-b322-8adc64ac5459 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9a16946-6a01-4984-8eba-7ea33532729f · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263505ab-b217-4b2e-b25a-28d22d546b0b · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160a6b59-3d94-4591-99d9-e8d9d0612bba · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Merging Improves Self-Critique Against Jailbreak Attacks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43554abd-4fb4-4256-9cc0-980f1fe26210 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb6fffd-c50d-49bc-81ba-9498562ad1d6 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs InAdvances in Neural Information Processing Systems (NeurIPS)
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10883db3-9c84-4ec1-bad8-edea276f3489 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Measuring Massive Multitask Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de3ed6b-2a59-4fd7-9736-82895f6c5e3e · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bca837-2cde-42f6-8e2d-0e43a7d6a0a2 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Editing Models with Task Arithmetic
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d59b9b3-4460-4231-8a1c-bdae26f7acce · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed595a6f-298f-4b54-97b1-b072efecf5e5 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f062de73-6dba-4e4b-b2f7-661c6350bfc7 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Ma, Q.; Liu, D.; Chen, Q.; Zhang, L.; and Shao, J
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34fe08dc-c22f-40d1-b7a6-97b24bf6ac3d · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e8fea2-8046-4f65-b74d-eb2de5219de0 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a8c56f-75eb-4375-ae08-1c6007877dbb · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6326a994-2fc5-488f-ba26-9a40adc8c0bb · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 661d7a26-7acf-4570-b022-dc295f2eff51 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e620d36b-1e62-46d8-af9a-40cd00a1cf4b · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Thakkar, M.; Fournier, Q.; Riemer, M.; Chen, P.-Y.; Zouaq, A.; Das, P.; and Chandar, S
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aed9cd0-699d-49f4-a2b1-3cbe23fd347a · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4dee33-b099-436b-821c-ee82e644ba1c · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0068464-d3d9-4cf4-9eef-4d141eb2f968 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs InIEEE Symposium on Security and Privacy (S&P)
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e8a0e393-1d01-47bf-9d95-30e914561159 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48faf8a9-8c10-4345-8027-8a2ed3018bd4 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs TIES-Merging: Resolving Interference When Merging Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc05a525-f187-4a51-91da-fc87494fd375 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295ba9dd-4fd7-4a7e-b69e-5d522a441f8d · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac8c5ef-3832-406b-b9f0-0167555c2ab6 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs 13 Zhang, J.; He, Y.; Cai, K.; Zhao, H.; Suya, F.; and Tian, Y
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 816b8940-0521-4aa9-8f69-e4b1d337326e · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs RogueMerge: Robust and Unified Attacks against LLM Model Merging
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98f5e297-da19-442b-ad61-891ff6f9b58f · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Instruction-Following Evaluation for Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c633179-6ea7-4e8f-95ce-bdbd5f951d6a · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5977ead-c045-42d7-8b35-919c91173192 · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs On Adaptive Attacks to Adversarial Example Defenses
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5912dd9-fde1-4e9a-ae1b-df767c20100c · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Evaluating Large Language Models Trained on Code
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a463a7-77b2-4c50-ba3f-5b93303b948c · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47639529-db75-489e-bd56-2232d42798dc · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44097e4-6445-4aeb-bbc8-f3be439fbf4f · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Refusal in Language Models Is Mediated by a Single Direction
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51554e35-6a93-4cf9-9a88-4642ef2b828a · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22524470-4f0d-4b99-b3eb-a6973346ebee · outbound
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Djuhera, A.; Kadhe, S
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.